Reliability Engineer 181585

PeopleSERVE, Inc.

  • Boston, TX
  • 30+ days ago

    Highlights

    We are looking for a systems thinking, reliability engineer who has helped teams scale through production insight, data and backup recovery, operational automation, developer guidance, real-time metrics, automation, automation, automation. Experience managing and interpreting large datasets using query languages and visualization tools(PowerBI/tableau), Experience in software development with Python, NodeJS, or Java with a focus on SDLC and automation.

    Numbers & Facts

    LocationBoston, TX

    Description

    We are looking for a systems thinking, reliability engineer who has helped teams scale through production insight, data and backup recovery, operational automation, developer guidance, real-time metrics, automation, automation, automation.

    • Strong background in several of the following: Go, Angular, Python, JavaScript, AWS, RESTful services, Ruby, MVC, Jenkins CI/CD, Configuration Automation (Chef, Ansible).
    • Preferred background in: Bootstrap, HTML/CSS, Shell Scripting, messaging frameworks (MQ), Service Oriented/Micro-service Architectures, OpenStack, Relational Databases (PostgreSQL).
    • Comfortable working in both Public and private cloud environments.
    • Crafting scalable solutions and automation to monitor the health and establish signals to drive understanding of our Container Platform environments.
    • Strengthening operational processes with Fidelity support and incident management teams for our cloud ecosystem
    • Working with Fidelity and cloud service provider product teams and driving ongoing reliability improvements in their Kubernetes service offerings.
    • Anticipating, discovering through ongoing interaction with, and prioritizing client / partner needs to serve as their voice and guide execution of the team.

    The Expertise You Have

    • Bachelor's Degree or equivalent experience in a technology related field (e.g. Computer Science, Engineering, etc.) required.
    • Production experience running Cloud and on-prem Storage workloads at scale
    • Experience managing and maintaining Kubernetes Clusters on EKS/AKS and RKS.
    • Demonstrates a drive for continuous improvement and enjoys tackling complex problems.
    • Experience managing and interpreting large datasets using query languages and visualization tools(PowerBI/tableau),
    • Experience in software development with Python, NodeJS, or Java with a focus on SDLC and automation
    • 5 -7 years of hands-on experience deploying and/or supporting highly distributed multi-tiered systems at scale.
    • Experience building and deploying Docker images including Docker Compose
    • Hands-on experience with Jenkins Core, including authoring and maintaining declarative CI/CD pipelines and libraries
    • Experience with distributed version control systems, Git preferred
    • Experience crafting and maintaining logging, monitoring, and alerting capabilities using tools like Datadog and Splunk
    • Practical experience in building cloud hosted and native applications for the enterprise. Maintains a deep understanding of a wide variety of AWS/Azure services that support reliability, observability, and automation/orchestration.
    • Experience in incident/crisis management and supporting critically important applications

    The Skills You Bring

    • Hands on experience with one or more observability tools (Prometheus, Grafana, ELK/OpenSearch, OpenTelemetry, Datadog, etc.)
    • Ability to automate with various scripting languages (Python, Shell scripting, etc.)
    • Experience managing systems using infrastructure as code tools (IAM, ARM, Terraform, Chef)

    Additional Value in Backup & Recovery:

    • Advance enterprise resiliency through improved recovery capabilities.
    • Reduce recovery time via automation.
    • Enable rehoused recovery into new datacenters.
    • Strengthen platform reliability through data protection design.

    Similar Jobs