Specialize in ensuring the reliability, resilience, and recoverability of enterprise services and systems across cloud, on-prem, and store environments. Serve as the SRE voice within incident and problem management, leading major incident response, driving root cause analysis, and partnering with Software Engineers and platform teams to reduce recurrence and improve system health in production.Qualifications3+ years of experience in Site Reliability Engineering, incident management, or production support for enterprise systemExperience leading or participating in major incident management (incident commander, bridge/war-room facilitation, executive communication during outagesExperience with problem management and root cause analysis (RCA), including driving corrective/preventive actions to closurSolid understanding of observability and monitoring concepts and tooling (e.g., Dynatrace, Azure Monitor) to detect, triage, and diagnose production issueWorking knowledge of Linux and scripting (e.g., BASH, Python) for troubleshooting and diagnostic automatioFamiliarity with Kubernetes, Docker, and cloud platforms (Azure, GCP) sufficient to troubleshoot and reason about system behavior in productioExperience working in an Agile environment, tracking work and metrics via a platform such as JirStrong communication skills — able to translate technical incident details for both engineering teams and business stakeholders under time pressur#J-18808-Ljbffr