Senior Site Reliability Engineer, Data & Analytics

Blizzard Entertainment

Irvine, CA(remote)

JOB DETAILS
SKILLS
Access Control, Accidental Death and Dismemberment (AD&D), Amazon Web Services (AWS), Apache Kafka, Automation, Budgeting, Cloud Computing, Communication Skills, Compensation and Benefits, Continuous Deployment/Delivery, Continuous Integration, Data Analysis, Disability Accommodations, Distributed Computing, Documentation, Fitness, Funding, GCP (Good Clinical Practices), GPU (Graphics Processing Unit), GitHub, Go Programming Language (Golang), Healthcare Reimbursement, Home Automation, Incident Response, Insurance, Jenkins, Legal, Linux Operating System, Load Testing, Machine Tool, Messaging Technology, Metrics, Operating Systems, Problem Solving Skills, Productivity Model, Python Programming/Scripting Language, Reliability Engineering, Rentals, Software Development
LOCATION
Irvine, CA
POSTED
Today

Senior Site Reliability Engineer, Data & AnalyticsThe Senior Site Reliability Engineer role is on our Data & Analytics team, partnering with data, analytics, ML, and platform engineering to improve the reliability, scalability, and performance of large‑scale data platforms, analytics pipelines, ML training pipelines, and inference services.In addition to core SRE responsibilities, this role will build operational and automation tooling that reduces toil, speeds up issue resolution, and improves engineering velocity. This includes contributing to internal platform services such as shared tooling, data integrations, and access‑control patterns used across Blizzard.The ideal candidate is a production‑minded SRE or platform engineer who is comfortable operating critical systems, writing software, and building tools that improve engineering efficiency without compromising reliability.This role is open to candidates based in Irvine, CA or Albany, NY (hybrid or on‑site), as well as fully remote candidates.ResponsibilitiesParticipate in an on‑call rotation and drive incidents to resolutionLead blameless postmortems and identify systemic reliability improvementsPartner with data, ML, and platform teams to improve batch, streaming, training, and inference workloadsSupport ML training pipelines and inference services, including GPU workloadsHelp define how data and ML services run on KubernetesDesign and build automation and operational tooling (e.g., workflows, diagnostic tooling, runbooks) to reduce on‑call burdenBuild and evolve centralized platform services, including shared tooling, data integrations, and access controlsDiagnose and resolve reliability, performance, and cost issues across distributed systemsChampion automation, documentation, and practices that reduce toilMaintain infrastructure using Terraform and infrastructure‑as‑code principlesImprove CI/CD and GitOps workflows (Jenkins, GitHub Actions, ArgoCD)Operate and improve containerized services on KubernetesDefine and measure reliability using SLIs, SLOs, and error budgetsRun load tests, capacity modeling, and production validationBuild internal tools and paved paths that help teams operate safely and efficientlyMinimum RequirementsExperience operating reliable, distributed systems in SRE, platform, or similar rolesExperience with data, analytics, ML, or large‑scale distributed workloadsStrong knowledge of Linux, containers, Kubernetes, and cloud infrastructureExperience building automation or internal tools (Python, Go, shell, etc.)Experience with infrastructure‑as‑code (e.g., Terraform)Experience with CI/CD or GitOps systems (e.g., Jenkins, GitHub Actions, ArgoCD)Familiarity with observability (metrics, logs, traces, alerting, incident response)Solid understanding of SRE concepts (SLIs, SLOs, error budgets, postmortems)Experience using modern development and automation practices to improve reliability and efficiencyExperience building internal tooling, automation, or developer productivity systemsStrong communication skills with technical and cross‑functional partnersBonus PointsExperience with data and ML systems (training pipelines, model serving, GPU workloads)Experience with distributed systems and messaging (Kafka, Pub/Sub)Experience working in Kubernetes‑based environmentsFamiliarity with observability tools (Prometheus, Grafana)Experience operating systems in cloud environments (GCP, AWS)BenefitsMedical, dental, vision, health savings account or health reimbursement account, healthcare spending accounts, dependent care spending accounts, life and AD&D insurance, disability insurance401(k) with company match, tuition reimbursement, charitable donation matchingPaid holidays and vacation, paid sick time, floating holidays, compassion and bereavement leaves, parental leaveMental health & wellbeing programs, fitness programs, free and discounted games, and a variety of other voluntary benefit programs like supplemental life & disability, legal service, ID protection, rental insurance, and othersRelocation assistance if the company requires geographic mobilityIn the U.S., the standard base pay range for this role is $101,000.00 – $186,754.00 annually. Compensation is based on experience, performance, and location.We are an equal opportunity employer and value diversity at our company. We do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, gender identity, age, marital status, veteran status, or disability status, among other characteristics.#J-18808-Ljbffr

About the Company

B

Blizzard Entertainment