Site Reliability Engineer

S:23 Recruitment

  • Santa Clara, California
  • 28 days ago

    Highlights

    Working alongside platform engineers, software developers, and technical leaders, you'll play a key role in evolving cloud-native infrastructure, driving operational excellence, and building the tooling that enables engineering teams to move quickly and safely. This role is focused on improving platform reliability, automation, observability, and developer experience while supporting enterprise-grade environments with strong security and compliance requirements.

    Numbers & Facts

    LocationSanta Clara, California

    Description

    S23 - Telecommunications

    Role - Site Reliability Engineer

    Location:Santa Clara, CA or Wall Township, NJ (Hybrid)

    Salary:$150,000 + Bonus + Comprehensive Benefits


    About the Role

    Join a team building and operating a modern Kubernetes-based platform that powers highly scalable, secure, and reliable cloud infrastructure. This role is focused on improving platform reliability, automation, observability, and developer experience while supporting enterprise-grade environments with strong security and compliance requirements.

    Working alongside platform engineers, software developers, and technical leaders, you'll play a key role in evolving cloud-native infrastructure, driving operational excellence, and building the tooling that enables engineering teams to move quickly and safely.

    Key Responsibilities

    • Own and operate core Kubernetes platform components, including deployments, upgrades, and ongoing maintenance
    • Design and implement scalable, resilient platform capabilities that improve reliability and developer productivity
    • Build automation, tooling, and CI/CD pipelines to streamline deployments and reduce operational overhead
    • Monitor platform health, troubleshoot production issues, and participate in on-call rotations and incident response
    • Lead post-incident reviews and implement preventative improvements to increase platform resilience
    • Define, measure, and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets
    • Support security and compliance initiatives through system hardening, documentation, and audit readiness
    • Partner with engineering teams to continuously improve platform performance, reliability, and operational processes

    What We're Looking For

    • 6+ years of experience in Site Reliability Engineering, DevOps, or Platform Engineering
    • Strong hands-on experience managing Kubernetes in production environments
    • Experience working with public cloud platforms such as AWS, Azure, or Google Cloud
    • Strong understanding of Linux administration, networking, and distributed systems
    • Experience with Infrastructure as Code tools such as Terraform
    • Proficiency in scripting or software development using Python, Go, Bash, or similar languages
    • Experience with monitoring and observability platforms including Prometheus, Grafana, and centralized logging solutions
    • Experience building and maintaining CI/CD pipelines and deployment automation
    • Strong troubleshooting skills with a focus on reliability, automation, and continuous improvement
    • Excellent collaboration and communication skills, with experience working across engineering teams



    Similar Jobs

    See more jobs