Apple Inc logo

Site Reliability Engineer, Apple Data Platform - AI/ML Platform

Apple Inc

  • Austin, TX
  • 5 days ago

    Highlights

    As an SRE on Apple Data Platform, youll operate and support the teams full portfolio, from big data pipelines to multi-cloud infrastructure, and grow into the teams go-to expert for ML/AI platform services - including Ray training and serving, LangGraph agent deployments, RAG architectures, embeddings platforms, and vector store platforms. We sit at the intersection of infrastructure, automation, and customer success - running incident response, providing hands-on support to internal teams, and partnering with developers to make cutting-edge services like Spark, Flink, Airflow, Ray, Notebooks and LLM-based agent platforms reliable at scale.

    Numbers & Facts

    LocationAustin, TX
    IndustryComputer/IT Services
    Company Size10,000 employees or more
    Year Founded1976
    Websitehttps://www.apple.com/jobs

    Description

    The Apple Services Engineering team (ASE) is one of the most exciting examples of Apples long-held passion for combining art and technology. These are the people who power the App Store, Apple TV, Apple Music, Apple Podcasts, and Apple Books - at extensive scale, meeting high expectations to deliver a huge variety of entertainment in over 35 languages to more than 150 countries.

    Within ASE, the Apple Data Platform SRE team keeps a massive, multi-cloud platform running for thousands of internal engineers building the next generation of data and AI products at Apple. We sit at the intersection of infrastructure, automation, and customer success - running incident response, providing hands-on support to internal teams, and partnering with developers to make cutting-edge services like Spark, Flink, Airflow, Ray, Notebooks and LLM-based agent platforms reliable at scale.

    This is a rare opportunity to build deep expertise across one of the most technically diverse platforms at Apple - while specialising in an area thats shaping the future of how Apple builds and operates AI. As an SRE on Apple Data Platform, youll operate and support the teams full portfolio, from big data pipelines to multi-cloud infrastructure, and grow into the teams go-to expert for ML/AI platform services - including Ray training and serving, LangGraph agent deployments, RAG architectures, embeddings platforms, and vector store platforms. You wont be building the models yourself, but youll be the infrastructure backbone behind the teams who do - keeping their services, pipelines, and platforms running flawlessly in production so they can focus on innovation. Were looking for a self-motivated engineer who thrives on ownership - someone who wants a set of services to call their own, the autonomy to drive their reliability roadmap, and the collaborative instinct to keep that work aligned with the teams broader direction. If you love solving hard operational problems, enjoy being the trusted expert customers turn to, and want a front-row seat to Apples ML/AI infrastructure evolution, this role offers real room to grow your scope and impact over time.Operate, monitor, and triage production and non-production environments across the ADP portfolio - data processing, ML/AI, and multi-cloud infrastructure. Participate in a rotating on-call schedule across supported services, including occasional weekday and weekend coverage. Own the operational health of ML/AI platform services as SME - driving reliability, support, and customer guidance for Ray, LangGraph agents, RAG pipelines, embeddings, and vector store platforms. Provide Slack-based support to internal customers; screen, triage, and resolve service related issues. Partner with dev teams across time zones to onboard new services - understanding architecture, then designing monitoring, alerting, and dashboards (Prometheus, Grafana, Splunk). Build automation and self-healing tooling that reduces manual toil and scales the teams operational capacity. Identify, escalate, and resolve production issues to protect platform reliability and customer experience. Collaborate with SRE and dev partner teams, engineering, and program management to align execution with team and org goals. Minimum Qualifications

    • Bachelors Degree in Computer Science, an engineering-related field, or equivalent related experience. 1-4 years in a Site Reliability Engineering, DevOps, or Infrastructure-focused role. Proficient in Python; working knowledge of Golang a plus. Experience with Kubernetes and at least one major cloud provider (AWS or GCP). Exposure to operating or supporting ML pipelines, model-serving infrastructure, or LLM-based systems in production. Strong communication skills and composure under pressure during incidents. Solid grounding in SRE principles, with prior on-call or production-support experience. Hands-on experience operating or supporting Ray (training/serving), LangGraph or similar agent orchestration frameworks, RAG architectures, embeddings platforms, or vector store platforms. Familiarity with MCP-based tooling and ML lifecycle/dataset management systems. Experience with S3 and cloud storage/networking fundamentals. Familiarity with observability tooling: Prometheus, Grafana, Splunk, PagerDuty. Working knowledge of CI/CD pipelines and deployment workflows. Deep understanding of one or more Big Data technologies (Spark, Flink, Airflow, Trino, Notebooks). A track record of automating manual operations through scripting or tooling. Intellectual curiosity and a drive to keep learning - for yourself, your team, and the org.

    About Company

    We bring amazing people together to make amazing things happen.

    We’re a diverse collection of thinkers and doers, continually reimagining what’s possible to help us all do what we love in new ways. The people who work here have reinvented entire industries with the Mac, iPhone, iPad, and Apple Watch, as well as with services, including iTunes, the App Store, Apple Music, and Apple Pay. And the same passion for innovation that goes into our products also applies to our practices — strengthening our commitment to leave the world better than we found it.

    About Apple

    There’s a place here for every kind of brilliant. Everyone here is an innovator, or an innovator-to-be, no matter what your team or your role. So bring your passion, courage, and original thinking and get ready to share it, because every new product, service, or feature we invent is the result of people working together to make each others’ ideas stronger. Innovation at this level depends on people who represent the variety of the human experience and inspire us with their own fresh perspectives. Together, we’ll do amazing work that can make a difference in people’s lives. Including your own. Learn more about working at Apple.

    Similar Jobs

    See more jobs