S23 - Telecommunications
Role - Site Reliability Engineer
Location:Santa Clara, CA or Wall Township, NJ (Hybrid)
Salary:$150,000 + Bonus + Comprehensive Benefits
About the Role
Join a team building and operating a modern Kubernetes-based platform that powers highly scalable, secure, and reliable cloud infrastructure. This role is focused on improving platform reliability, automation, observability, and developer experience while supporting enterprise-grade environments with strong security and compliance requirements.
Working alongside platform engineers, software developers, and technical leaders, you'll play a key role in evolving cloud-native infrastructure, driving operational excellence, and building the tooling that enables engineering teams to move quickly and safely.
Key Responsibilities
- Own and operate core Kubernetes platform components, including deployments, upgrades, and ongoing maintenance
- Design and implement scalable, resilient platform capabilities that improve reliability and developer productivity
- Build automation, tooling, and CI/CD pipelines to streamline deployments and reduce operational overhead
- Monitor platform health, troubleshoot production issues, and participate in on-call rotations and incident response
- Lead post-incident reviews and implement preventative improvements to increase platform resilience
- Define, measure, and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets
- Support security and compliance initiatives through system hardening, documentation, and audit readiness
- Partner with engineering teams to continuously improve platform performance, reliability, and operational processes
What We're Looking For
- 6+ years of experience in Site Reliability Engineering, DevOps, or Platform Engineering
- Strong hands-on experience managing Kubernetes in production environments
- Experience working with public cloud platforms such as AWS, Azure, or Google Cloud
- Strong understanding of Linux administration, networking, and distributed systems
- Experience with Infrastructure as Code tools such as Terraform
- Proficiency in scripting or software development using Python, Go, Bash, or similar languages
- Experience with monitoring and observability platforms including Prometheus, Grafana, and centralized logging solutions
- Experience building and maintaining CI/CD pipelines and deployment automation
- Strong troubleshooting skills with a focus on reliability, automation, and continuous improvement
- Excellent collaboration and communication skills, with experience working across engineering teams