Sr. Site Reliability Engineer, Data Center & Network Infrastructure

Tesla Inc

  • Austin, TX
  • 22 days ago

    Highlights

    This role bridges network operations and data center facility infrastructure-including bare-metal server reliability, power/cooling systems, and network fabric-with a relentless focus on uptime, scalability, and user experience. Lead Incident Response: Participate in high-severity infrastructure on-call rotations, directing rapid triage, mitigation, and root cause analysis (RCA) for server, network, and facility anomalies.

    Numbers & Facts

    LocationAustin, TX

    Description

    The Enterprise Infrastructure SRE team is seeking a seasoned Sr. Site Reliability Engineer to own the reliability, automation, and observability of our data center and network infrastructure. This role bridges network operations and data center facility infrastructure-including bare-metal server reliability, power/cooling systems, and network fabric-with a relentless focus on uptime, scalability, and user experience.

    You will minimize manual toil through advanced automation, drive operational excellence through blameless postmortems, and proactively mitigate infrastructure risks. As a senior engineer, you will lead high-impact projects, mentor peers by example, and foster continuous improvement across the global SRE organization.

    • Enhance Telemetry Platforms: Scale observability, logging, and alerting systems using Grafana, Splunk, and Prometheus

    • Build Correlated Dashboards: Design visualization tools that connect server health, network telemetry, and facility power/cooling performance across fragmented data sources

    • Drive Data Insights: Write complex SQL and SPL queries to analyze infrastructure trends, isolate production bottlenecks, and surface environmental health insights via IPMI interfaces

    • Eliminate Manual Toil: Develop robust automation scripts and tooling to handle hardware incident triage, alert noise reduction, and log correlation

    • Manage Source of Truth: Maintain and scale Netbox inventory systems, building automated API pipelines to track physical layout, device lifecycles, and cable topologies

    • Standardize Operational Playbooks: Create and maintain high-quality runbooks, KB articles, and SOPs to enable bot-assisted incident resolution

    • Lead Incident Response: Participate in high-severity infrastructure on-call rotations, directing rapid triage, mitigation, and root cause analysis (RCA) for server, network, and facility anomalies

    • Cross-Functional Partnership: Collaborate with architecture, deployment, hardware engineering, and facility operations teams to ensure new implementations are supportable and monitored from day one

    • Optimize Performance: Continuous monitoring of network and infrastructure performance to execute sustainability and optimization changes

    • Bachelor s Degree in Computer Science, Electrical Engineering, or a related technical field, or equivalent practical experience

    • 8+ years of experience in Site Reliability Engineering (SRE), network operations, or data center infrastructure management across highly distributed, large-scale environments

    • Strong proficiency in Python, Go, or Shell scripting, combined with configuration management frameworks like Salt or Ansible

    • Observability Expertise: Deep production experience with Prometheus, Grafana, Alert Manager, and Splunk for enterprise metrics and logging

    • Strong SQL skills (PostgreSQL, MySQL) and a proven track record of consuming and building RESTful APIs to integrate infrastructure tooling

    • Advanced Linux system fundamentals paired with hands-on experience provisioning, troubleshooting, and managing bare-metal enterprise server architectures

    • Deep understanding of TCP/UDP, IPv4/IPv6, BGP, EVPN, VxLAN, Segment Routing, and load balancing

    • Experience managing enterprise hardware vendors like Arista Networks, Juniper Networks, Cisco, Palo Alto Networks firewalls, and F5 load balancers

    • Practical knowledge of out-of-band management (IPMI), PDU architecture, and data center physical cooling systems (liquid cooling, HVAC, hot/cold aisle containment)

    Benefits

    Along with competitive pay, as a full-time Tesla employee, you are eligible for the following benefits at day 1 of hire:

    • Medical plans > plan options with $0 payroll deduction
    • Family-building, fertility, adoption and surrogacy benefits
    • Dental (including orthodontic coverage) and vision plans, both have options with a $0 paycheck contribution
    • Company Paid (Health Savings Accounts) HSA Contribution when enrolled in the High-Deductible medical plan with HSA
    • Healthcare and Dependent Care Flexible Spending Accounts (FSA)
    • 401(k) with employer match, Employee Stock Purchase Plans, and other financial benefits
    • Company paid Basic Life, AD&D
    • Short-term and long-term disability insurance (90 day waiting period)
    • Employee Assistance Program
    • Sick and Vacation time (Flex time for salary positions, Accrued hours for Hourly positions), and Paid Holidays
    • Back-up childcare and parenting support resources
    • Voluntary benefits to include: critical illness, hospital indemnity, accident insurance, theft & legal services, and pet insurance
    • Weight Loss and Tobacco Cessation Programs
    • Tesla Babies program
    • Commuter benefits
    • Employee discounts and perks program

    Similar Jobs

    See more jobs