Site Reliability Engineer (SRE) Production Services

TechDigital Corporation

  • Pittsburgh, PA
  • 30+ days ago

    Highlights

    Build reliability dashboards tracking incidents, repeat issues, failure rates, and automation coverage. Define and manage Service Level Objectives (SLOs) for critical services and batch processes.

    Numbers & Facts

    LocationPittsburgh, PA
    IndustryOther/Not Classified
    Company Size100 to 499 employees

    Description

    Mandatory Skills:
    1. Java Spring boot 2. Apache Kafka 3. Dev Ops 4. CI/CD automation

    Years of experience required: 8-10

    Job Description:
    Automation & Efficiency
    Automate the top 5 high-volume support and request types
    Build self-service and agent-driven solutions to reduce manual work
    Harden operational workflows for consistency, auditability, and resilience
    Implement auto-retry and backoff for recurring failure patterns

    Reliability Engineering
    Define and manage Service Level Objectives (SLOs) for critical services and batch processes
    Apply error budget concepts to guide reliability and release decisions
    Improve batch reliability through standardized recovery patterns and monitoring

    Observability & Metrics
    Build reliability dashboards tracking incidents, repeat issues, failure rates, and automation coverage
    Improve operational reporting and visibility across incidents, problems, and changes

    Runbooks & Self-Service
    Develop and expand runbooks for key production scenarios
    Convert runbooks into automated remediation workflows
    Enable self-service for repeat operational requests
    Drive conversion of repeat incidents into permanent fixes and known problems

    Self-Healing & Intelligent Operations
    Implement self-healing capabilities to minimize manual intervention
    Optimize alerting systems (e.g., Moogsoft) to reduce noise and improve signal quality
    Leverage automation and AI to resolve recurring issues with minimal human involvement

    Similar Jobs