5+ years of engineering management experience leading SRE, infrastructure, or observability/monitoring teams Experience hiring and leading engineers, with a desire to build, grow, and mentor a team Deep understanding of observability systems and practices: metrics, logging, tracing, alerting, SLOs, error budgets, and fault analysis at scale Strong systems background, comfortable troubleshooting across the full stack (network, OS, container runtime, application) Experience operating large-scale, multi-tenant distributed systems in production, including Kubernetes environments Practical, solid knowledge of shell/bash scripting and at least one higher-level production language (Python preferred; Go, Java, or Scala also valued) Demonstrated experience applying AI/ML tooling or LLM-based solutions to improve SRE or infrastructure operations Track record of building high-performing teams through coaching, clear expectations, and psychological safety Demonstrated ability to drive cross-functional initiatives to completion and communicate at the executive level Bachelors or Masters degree in Computer Science, Engineering, or related field, or equivalent experienceDeep familiarity with the Prometheus ecosystem and cloud-native observability stacks (Thanos, Splunk, OpenTelemetry, or similar) Experience with third-party cloud platforms (AWS, GCP, or Azure) and infrastructure as code (Terraform, Ansible) Comfortable with open-source configuration management and orchestration tools (Helm, Puppet, Spinnaker) Demonstrable knowledge of TCP/IP, HTTP, web application security, and multi-tier web application architectures Experience running infrastructure as an internal managed service with defined SLAs Familiarity with microservices architecture and container orchestration with Kubernetes at scale Background in capacity planning, performance engineering, or infrastructure architecture Track record of driving cultural and process transformation within SRE organizations Experience building or deploying AI-powered operational tooling (AIOps, intelligent alerting, automated diagnostics) Developing and delivering multi-mode communications tailored to the unique needs of different audiences Anticipating and balancing the needs of multiple stakeholders Making sense of complex, high-quantity, and sometimes contradictory information to solve problems effectively Rebounding from setbacks and adversity when facing difficult situations Knowing the most effective and efficient processes to get things done, with a focus on continuous improvement. Lead, grow, and mentor a team of Site Reliability Engineers, conducting regular 1:1s, performance reviews, and career development discussions Hire and build out the SRE team, with a desire to develop engineers to meet both their career goals and the organizations goals Build and scale a high-performing team through coaching, clear expectations, and psychological safety Foster an inclusive team culture and mentor diverse talent Champion engineering best practices for code quality, system design, and operational excellence across the broader organization.