Apple Inc logo

AIML - Distinguished Engineer, Foundation Model

Apple Inc

  • Cupertino, CA
  • 5 days ago

    Highlights

    Architect inference systems that support a wide range of internal use cases - data generation and rollouts for training, offline and large-scale evaluation, and judge/reward scoring - across text, image, speech, and multi-modal models, each with distinct throughput, cost, and quality constraints. You will partner closely with modeling and research teams to bring new capabilities into the development loop, work across many internal teams with very different requirements, and lead a diverse set of engineers in turning an ambitious vision into shipped milestones.

    Numbers & Facts

    LocationCupertino, CA
    IndustryComputer/IT Services
    Company Size10,000 employees or more
    Year Founded1976
    Websitehttps://www.apple.com/jobs

    Description

    Apple is revolutionizing artificial intelligence by developing sophisticated foundation models that power intelligent features across our product ecosystem. We are seeking a Distinguished Engineer to set the technical direction for the systems that power our foundation model training - with an initial focus on the inference engine that the foundation model team relies on for training, model evaluation, and other needs in the model development loop.

    This is a senior individual-contributor leadership role. You will be one of the most senior technical voices for foundation model systems at Apple: defining the vision, driving execution across many teams, and raising the bar for engineering excellence. The role starts with the inference engine, but we expect you to move fluidly into adjacent training systems areas as the needs of the foundation model program evolve. engineered specifically for Apple silicon and for experiences that are private, personal, and deeply integrated into the OS. Behind that modeling work sits a demanding systems layer, and the inference engine is at its center.

    Our inference engine is used by the foundation model team throughout the model development lifecycle: generating and processing data and running rollouts for training, powering large-scale model evaluation, and serving as an LLM judge that scores and compares model outputs. These workloads are high throughput, bursty, and tightly coupled to research iteration - the speed, efficiency, and reliability of the engine directly set the pace at which the team can train and improve models.

    As a Distinguished Engineer, you will own the technical strategy for this inference engine and the broader systems that support it. You will partner closely with modeling and research teams to bring new capabilities into the development loop, work across many internal teams with very different requirements, and lead a diverse set of engineers in turning an ambitious vision into shipped milestones. While inference is the initial focus, you will also help shape adjacent areas - training infrastructure, data systems, and evaluation. If you are drawn to hard systems problems where the research and the infrastructure are inseparable, this is the role.Set and drive the technical vision and roadmap for the foundation model teams inference engine and the systems around it, used for training, evaluation, and LLM-as-judge workloads. Lead deep work on inference performance, efficiency, and reliability: throughput and latency optimization, batching and scheduling, quantization, speculative decoding, KV-cache management, memory and compute efficiency, and hardware-aware optimization. Architect inference systems that support a wide range of internal use cases - data generation and rollouts for training, offline and large-scale evaluation, and judge/reward scoring - across text, image, speech, and multi-modal models, each with distinct throughput, cost, and quality constraints. Extend your impact into adjacent systems areas - training infrastructure, data pipelines, and evaluation harnesses. Partner with many teams that depend on the engine, translating their diverse needs into a coherent platform, clear interfaces, and a prioritized roadmap. Work closely with ML researchers and modeling teams to co-design models and systems, and to bring state-of-the-art techniques from prototype into the development loop reliably. Lead a diverse set of engineers across teams in setting direction and executing against it; align stakeholders, resolve technical trade-offs, and make the calls that keep large efforts moving. Drive prioritization and milestone delivery across competing demands, balancing near-term research needs against long-term platform investment. Mentor and grow junior and senior engineers; establish engineering standards, review designs, and multiply the impact of the organization.MS or PhD in Computer Science, Machine Learning, or related technical field, or equivalent industry experience. 15+ years of experience building large-scale ML or distributed systems, with a track record of technical leadership and industry-wide or company-wide impact. Deep, hands-on expertise in foundation model inference engines, with a proven record of improving performance, efficiency, and reliability at scale. Deep experience supporting a diverse set of foundation model inference use cases, each with different throughput, latency, cost, and quality constraints. Breadth beyond inference - the ability to contribute in adjacent systems areas such as training infrastructure, data systems, or evaluation. Deep understanding of GPU/TPU/accelerator architecture, distributed systems, and model optimization (quantization, distillation, compilation, serving). Proficiency with ML frameworks such as JAX, PyTorch, and with inference/serving stacks. Proven experience leading a diverse set of engineers in setting vision and driving execution, including prioritization for milestone deliveries. Demonstrated experience mentoring junior and senior engineers. Demonstrated experience partnering with ML researchers and modeling teams to productionize research.Experience building or leading inference systems for large language models and multi-modal foundation models at scale. Experience with inference in training, evaluation, or reinforcement-learning loops (e.g., large-scale rollouts, offline eval, or LLM-as-judge / reward scoring). Familiarity with Kubernetes, Docker, and cloud platforms (AWS, GCP, Azure), and with distributed computing frameworks. History of defining technical strategy that shaped an organizations or the industrys direction.

    About Company

    We bring amazing people together to make amazing things happen.

    We’re a diverse collection of thinkers and doers, continually reimagining what’s possible to help us all do what we love in new ways. The people who work here have reinvented entire industries with the Mac, iPhone, iPad, and Apple Watch, as well as with services, including iTunes, the App Store, Apple Music, and Apple Pay. And the same passion for innovation that goes into our products also applies to our practices — strengthening our commitment to leave the world better than we found it.

    About Apple

    There’s a place here for every kind of brilliant. Everyone here is an innovator, or an innovator-to-be, no matter what your team or your role. So bring your passion, courage, and original thinking and get ready to share it, because every new product, service, or feature we invent is the result of people working together to make each others’ ideas stronger. Innovation at this level depends on people who represent the variety of the human experience and inspire us with their own fresh perspectives. Together, we’ll do amazing work that can make a difference in people’s lives. Including your own. Learn more about working at Apple.

    Similar Jobs

    See more jobs