Vice President of Site Reliability Engineering

Selby Jennings Ltd

  • Dallas, TX
  • 6 days ago
  • Remote

    Highlights

    Our client is seeking a Vice President of Site Reliability Engineering to lead infrastructure automation and platform reliability initiatives within a highly scalable hybrid environment. Experience implementing monitoring and observability solutions using tools such as Grafana, Prometheus, Splunk, ELK, or related platforms.

    Numbers & Facts

    LocationDallas, TX (
    Remote
    )

    Description

    About the Job

    Our client is seeking a Vice President of Site Reliability Engineering to lead infrastructure automation and platform reliability initiatives within a highly scalable hybrid environment. This individual will be responsible for driving automation strategy, establishing Infrastructure as Code standards, overseeing platform reliability, and mentoring a team of engineers focused on building and maintaining modern infrastructure services. The role offers the opportunity to remain hands-on while influencing the long-term direction of infrastructure engineering across the organization.

    What You'll Do:

    • Define and enforce Infrastructure as Code standards to drive consistent, scalable, and secure deployments.
    • Oversee configuration management, system provisioning, image lifecycle management, and infrastructure automation across Windows and Linux environments.
    • Develop and enhance self-service platforms that improve engineering efficiency, deployment velocity, and operational stability.
    • Own monitoring, observability, and reliability initiatives for internal infrastructure and automation platforms.
    • Partner with infrastructure, security, networking, and engineering teams to improve operational workflows and platform capabilities.
    • Drive the development of internal tooling and automation solutions using modern scripting and programming languages.
    • Establish best practices around infrastructure lifecycle management, performance optimization, and platform scalability.

    What We're Looking For:

    • 8+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or Infrastructure Engineering.
    • Deep expertise with Terraform and Infrastructure as Code practices in enterprise environments.
    • Strong hands-on experience with Ansible and large-scale automation initiatives.
    • Experience managing and automating VMware-based environments as well as public cloud platforms such as AWS and Azure.
    • Proficiency with scripting and automation using Python, PowerShell, Bash, Go, or similar languages.
    • Experience implementing monitoring and observability solutions using tools such as Grafana, Prometheus, Splunk, ELK, or related platforms.
    • Strong understanding of CI/CD pipelines, Git workflows, and modern software delivery practices.
    • Ability to support and troubleshoot both Windows and Linux infrastructure environments.

    Preferred Qualifications:

    • Previous experience leading or managing engineering teams.
    • Experience with identity and access management platforms.
    • Exposure to enterprise storage, backup, disaster recovery, and large-scale infrastructure environments.
    • Strong understanding of networking concepts, infrastructure architecture, and systems performance optimization.

    Compensation & Benefits:

    • Annual performance bonus
    • Equity participation
    • Fully remote work environment
    • Comprehensive medical, dental, and vision coverage
    • 401(k) with company contribution
    • Generous PTO and paid holidays
    • Paid parental leave
    • Opportunity to help shape and scale a modern infrastructure automation function within a growing technology-driven organization.

    Similar Jobs

    See more jobs