Must Have Technical/Functional Skills
7+ years of experience in SRE, platform engineering, or cloud infrastructure engineering in large-scale enterprise environments Deep, hands-on expertise with Microsoft Azure - minimum 4 years in a primary Azure cloud engineering role. Expert-level proficiency with AKS: cluster lifecycle management, RBAC, network policies, pod security standards, cluster autoscaler, and Workload Identity. Strong experience in Microservices development using Java and perform CI/CD using Azure DevOps (ADO) Experience designing and operating enterprise observability platforms using Dynatrace Demonstrable track record of owning SLOs/SLIs and delivering measurable reliability improvements in production.
Roles & Responsibilities
Define, own, and enforce enterprise-wide SLOs, SLIs, and Error Budgets across all Tier-0 and Tier-1 Azure-hosted services, report SLA compliance to executive stakeholders monthly. Lead architectural reviews for new services and ensure reliability non-functionals (availability targets, RTO/RPO) are embedded from design through to production. Champion and implement chaos engineering practices Drive Disaster Recovery (DR) design and conduct quarterly DR drills across Azure paired regions. Incident Management & On-Call Serve as Incident Commander for P1/P2 major incidents, own end-to-end incident lifecycle from detection through resolution and Post-Incident Review (PIR). Participate in a structured On-Call rotation with follow-the-sun global coverage; maintain response SLAs of