Systems Reliability Engineer

Pentangle Tech

  • Seattle, WA
  • 3 days ago

    Highlights

    Strong knowledge of application and infrastructure technologies monitored by Dynatrace, including Java, NET, Node.js, Python, JavaScript, Go, SQL databases, Linux, Windows Server, web servers, APIs, microservices, and Kubernetes-based cloud-native architectures. Proficiency integrating Dynatrace with cloud platforms (AWS, Azure, GCP), Kubernetes, Docker, OpenTelemetry, CI/CD pipelines (GitHub, Azure DevOps, Harness, Jenkins), ITSM platforms (ServiceNow), and collaboration tools (Slack, Microsoft Teams).

    Numbers & Facts

    LocationSeattle, WA

    Description

    Job Title: Systems Reliability Engineer
    Location: Seattle WA
    Duration: Long term
    Job Description:
    Role Summary
    The Systems Reliability Engineer (SRE) is a subject matter expert in software engineering and IT operations. As an individual contributor, this role exercise considerable judgement to ensure the resilience and stability of our applications and supporting infrastructure within an enterprise SRE center of excellence.
    Dynatrace specific requirements:
    • Expert-level experience administering and optimizing Dynatrace, including OneAgent deployment, environment configuration, management zones, tagging, dashboards, alerting, and synthetic monitoring.
    • Deep understanding of Dynatrace observability capabilities, including Distributed Tracing, PurePath , Real User Monitoring (RUM), Digital Experience Monitoring (DEM), Infrastructure Monitoring, Log Management, and Application Security.
    • Proficiency integrating Dynatrace with cloud platforms (AWS, Azure, GCP), Kubernetes, Docker, OpenTelemetry, CI/CD pipelines (GitHub, Azure DevOps, Harness, Jenkins), ITSM platforms (ServiceNow), and collaboration tools (Slack, Microsoft Teams).
    • Experience leveraging Dynatrace Query Language (DQL), Dynatrace API, Workflows, Grail, and Davis AI to automate monitoring, develop custom dashboards, perform advanced analytics, and accelerate incident detection and root cause analysis.
    • Strong knowledge of application and infrastructure technologies monitored by Dynatrace, including Java, .NET, Node.js, Python, JavaScript, Go, SQL databases, Linux, Windows Server, web servers, APIs, microservices, and Kubernetes-based cloud-native architectures.
    Key Duties
    • Apply deep knowledge to develop and maintain software tools to enhance system reliability.
    • Lead teams in building real-time monitoring and observability of applications to preemptively identify potential issues.
    • Manage the response to other business partners to minimize service disruptions and maintain high availability.
    • Collaborate with development teams to ensure seamless code deployment and operational excellence.
    • Engaging in system design consulting, platform management, and capacity planning to support scalability and performance goals.
    • Influence and guide other teams to improve systems support, monitoring, and administration.
    • Exercise considerable judgment to make decisions pertaining to operational compliance with all security, privacy, audit, disaster recovery, and other company requirements.
    • Support regulatory audit requirements.
    Job-Specific Skills, Experience & Education
    Required
      • 4 years of experience in information technology.
      • A Bachelor's degree, preferably with a focus in computer science, engineering, information systems, or an additional two years of training/experience in lieu of this degree.
      • Demonstrate experience in coaching and mentoring system engineers.
      • Minimum age of 18.
      • Must be authorized to work in the U.S.
      • High school diploma or equivalent is required.
    Preferred
      • Experience with SRE practices, agile methodologies, development lifecycles, and DevOps best practices.
      • Experience applying ITIL and IT process best practices.
      • Technical knowledge of application designs and architectures.
      • Experience with technical engineering working in IT operations.
      • Excellent communication skills and a proven ability to collaborate with a variety of team members.
      • Ability to work collaboratively with cross functional teams to understand objectives, gather automation requirements, write technical specifications and perform in a lead role.
      • Experience with event correlation.
      • Experience with Windows and Linux based operating environments.
      • Experience with multi-cloud environments.
      • Proven ability to successfully work with multiple vendors.
      • Strong interpersonal, organizational, communication, and customer service skills.
    Job-Specific Leadership Expectations
    • Embody our values to own safety, do the right thing, be kind-hearted, deliver performance, and be remarkable

    Similar Jobs

    See more jobs