Champion AI-powered tooling and automation to improve incident triage, reduce operational toil, drive capacity efficiency, and accelerate engineering workflows Lead, mentor, and grow a team of Software and SRE engineers across multiple geographies, fostering a culture of ownership, collaboration, and continuous improvement Establish and maintain a sustainable 24/7 on-call rotation with clear escalation paths, severity definitions, and response time SLAs across US and UK locations Oversee release engineering and deployment automation, including CI/CD pipelines, canary deployments, and zero-downtime rollout strategies Manage infrastructure modernization initiatives including Kubernetes control plane operations, database migrations, and configuration management evolution Drive incident management excellence - including post-incident reviews, preventive measures, and production readiness reviews for all releases5+ years of experience managing infrastructure, SRE, or platform engineering teams operating large-scale distributed systems Proven track record of building and leading on-call organizations with structured incident management, escalation procedures, and post-incident review processes Strong technical background in cloud infrastructure, compute orchestration, and bare metal provisioning at scale Experience with Kubernetes, OpenStack, KVM/hypervisor technologies, and Infrastructure as Code tools (Chef, Ansible, Terraform, or Salt) Deep understanding of SRE principles including SLOs, error budgets, capacity planning, and release engineering Excellent verbal and written communication skills with the ability to influence across teams and levels Demonstrated ability to recruit, develop, and retain high-performing engineering talentHands-on experience leveraging AI and machine learning to improve operational efficiency, incident management, or infrastructure automation Experience managing or scaling batch compute, job scheduling, or HPC platforms Proficiency in Go or Python with a strong automation-first mindset Familiarity with observability stacks (Prometheus, Grafana, distributed tracing) and centralized logging at scale Experience operating large-scale multi-tenant Infrastructure as a Managed Service Experience managing geographically distributed teams and follow-the-sun on-call models Track record of driving capacity efficiency initiatives resulting in measurable cost optimization. Apple Service Engineering (ASE)s Compute team is seeking an experienced Software Engineering Manager to lead a team of Infrastructure and Site Reliability Engineers responsible for operating and scaling large-scale batch compute infrastructure across Apples data centers.