Apolis logo

Site Reliability Engineer (SRE) Production Services

Apolis

  • Pittsburgh, PA
  • 30+ days ago
  • $50–$60 Per Hour

Highlights

We are seeking an experienced Site Reliability Engineer (SRE) Production Services to support and enhance production operations through automation, reliability engineering, observability, and self-healing capabilities. The ideal candidate will have strong expertise in Java Spring Boot, Apache Kafka, DevOps, and CI/CD automation , along with experience building scalable, resilient, and highly available production systems.

Numbers & Facts

LocationPittsburgh, PA
IndustryComputer/IT Services
Salary$50–$60 Per Hour
Company Size500 to 999 employees
Websitehttps://www.apolisrises.com/

Description

Job Title: Site Reliability Engineer (SRE) Production Services

Location: Pittsburgh, PA 15219 (Onsite)/Local Candidates Only

Tax Term (W2, C2C): W2

Job Type (Permanent/Contract): Contract

Duration: Long Term

Description:

We are seeking an experienced Site Reliability Engineer (SRE) Production Services to support and enhance production operations through automation, reliability engineering, observability, and self-healing capabilities. The ideal candidate will have strong expertise in Java Spring Boot, Apache Kafka, DevOps, and CI/CD automation, along with experience building scalable, resilient, and highly available production systems. This is a fully onsite role in Pittsburgh, PA, and only local candidates will be considered.

Role and Responsibilities:

  • Automate high-volume production support requests and operational workflows.
  • Develop self-service and agent-driven automation solutions to minimize manual effort.
  • Implement standardized operational processes with auditability and resilience.
  • Build auto-retry, backoff, and recovery mechanisms for recurring production failures.
  • Define, monitor, and maintain Service Level Objectives (SLOs) and apply error budget principles.
  • Improve reliability of batch processing through standardized recovery patterns.
  • Develop observability dashboards for incidents, failures, automation coverage, and operational metrics.
  • Create and enhance production runbooks and convert them into automated remediation workflows.
  • Drive permanent resolution of recurring production issues through root cause analysis.
  • Implement self-healing capabilities to reduce operational intervention.
  • Optimize monitoring and alerting platforms (Moogsoft or similar) to improve signal-to-noise ratio.
  • Leverage automation and AI-driven operational solutions for recurring production issues.
  • Collaborate with development, infrastructure, and operations teams to improve system reliability and production stability.

Required Skills:

  • 12+ years of overall IT experience.
  • 8 10+ years of Site Reliability Engineering (SRE) or Production Support experience.
  • Strong hands-on experience with Java and Spring Boot.
  • Experience with Apache Kafka.
  • Strong knowledge of DevOps practices and tools.
  • Expertise in CI/CD automation (Jenkins, GitLab CI, Azure DevOps, etc.).
  • Experience with production monitoring, observability, dashboards, and alerting tools.
  • Knowledge of Service Level Objectives (SLOs), SLIs, and Error Budgets.
  • Experience implementing automation, self-healing, and operational runbooks.
  • Strong troubleshooting and root cause analysis skills.
  • Experience working in enterprise production support environments.

Qualifications:

  • Bachelor's degree in Computer Science, Information Technology, Engineering, or related field.
  • Experience with cloud platforms and container technologies is a plus.
  • Excellent communication and collaboration skills.
  • Ability to work in a fast-paced production support environment.
  • Local candidates available to work onsite in Pittsburgh, PA.

Benefits

Paid Sick Days, Employee Referral Program, Employee Events, Retirement / Pension Plans

About Company

Since 1996, RJT has provided successful SAP, Oracle, and IT consulting solutions and staffing services to clients around the world. The new Apolis brings you the same personalized service fortified with a greater array of IT solutions, global expertise, and cost-management strategies.

We are a global IT consultancy that seamlessly integrates experts and leading-edge solutions into your organization so you can focus on what really matters.

Similar Jobs