Apple Inc logo

AI Engineer - Algorithm Evaluation & Agentic Systems

Apple Inc

  • Sunnyvale, CA
  • 2 days ago

    Highlights

    What We Value Production mindset: correctness, observability and maintainability Ability to reason about system-level tradeoffs, not just model performance Ability to balance experimentation speed with engineering rigor Comfort working in ambiguous problem spaces and defining metrics from first principles Clear communication of technical findings to both technical and non-technical audiences Within the DAQ team, our core mission is to evaluate and elevate advanced visual technologies. Agentic Architecture: Build, deploy, and evaluate agentic workflows that utilize these vision models to autonomously solve multi-step user problems (e.g., video summarization, visual search).

    Numbers & Facts

    LocationSunnyvale, CA
    IndustryComputer/IT Services
    Company Size10,000 employees or more
    Year Founded1976
    Websitehttps://www.apple.com/jobs

    Description

    How do we ensure Apples next-generation AI products are robust, safe, and truly intelligent? Join the DAQ team to help answer that. We are seeking an AI Engineer specializing in algorithm evaluation and agentic systems design for advanced computer vision and video understanding algorithms.

    What We Value Production mindset: correctness, observability and maintainability Ability to reason about system-level tradeoffs, not just model performance Ability to balance experimentation speed with engineering rigor Comfort working in ambiguous problem spaces and defining metrics from first principles Clear communication of technical findings to both technical and non-technical audiences Within the DAQ team, our core mission is to evaluate and elevate advanced visual technologies. As a key member of this group, you will lead the benchmarking and integration of state-of-the-art models for image and video understanding. Rather than focusing on core model training, you will apply your deep CV and ML expertise to rigorously test models in applied settings, uncover edge-case failure modes, and architect advanced agentic systems. If you are passionate about AI safety, robust evaluation, and building autonomous multi-modal workflows that bridge experimentation with production, we'd love to hear from you.Algorithm Evaluation & Benchmarking: Design, build, and scale comprehensive evaluation pipelines. You will be responsible for both holistic end-to-end system evaluation and granular component-level testing to rigorously measure model capabilities on complex image and video understanding tasks. Deep Failure Analysis: Leverage your CV and ML background to dive deep into model outputs, identifying root causes of visual hallucinations, temporal inconsistencies in video, and edge-case failures. Agentic Architecture: Build, deploy, and evaluate agentic workflows that utilize these vision models to autonomously solve multi-step user problems (e.g., video summarization, visual search). You will heavily utilize component-level evaluation to isolate and triage exactly which parts of the agentic workflow (e.g., tool selection, memory retrieval, visual reasoning) are succeeding or failing. Golden Data Curation: Lead the strategy for curating high-quality, schematized datasets and ground-truth benchmarks specifically tailored for evaluating multi-modal capabilities. Cross-Functional Collaboration: Partner closely with the core model training teams. You will provide them with actionable, data-driven insights and metrics to guide the next iteration of model training and fine-tuning.MS and a minimum of 3 years relevant industry experience 3+ years of applied experience in Machine Learning, Computer Vision, or AI System Evaluation Solid ML Foundation: Deep understanding of core Machine Learning principles, including probability, statistics, data distributions, and model bias/variance. You can apply statistical rigor to ensure evaluation metrics are meaningful and reliable. Computer Vision Expertise: Deep theoretical and practical understanding of Computer Vision (CV) and Vision-Language Models (VLMs). You must understand how Vision Transformers (ViTs), spatial-temporal modeling, and image/video processing work under the hood to effectively evaluate them. Advanced Evaluation Skills: Proven track record of defining robust metrics/KPIs and designing rigorous evaluation frameworks for generative AI or foundation models. Deep experience with custom benchmark creation, automated regression testing, LLM/VLM-as-a-judge methodologies, and human-in-the-loop evaluation. Agentic Systems: Experience building and evaluating LLM/VLM-powered agents, including tool use, multi-step reasoning, planning, and memory management workflows. Failure Analysis: Strong intuition for probing ML models to discover edge cases, hallucinations, and performance bottlenecks in constrained environments. Be able to translate findings into actionable improvement recommendations. Engineering Excellence: Strong proficiency in Python and experience with deep learning frameworks (PyTorch) for running inference, extracting embeddings, and building scalable evaluation pipelines.Demonstrated ability to lead technical evaluation strategies end-to-end, drive architectural decisions for testing infrastructure, and mentor engineers. Strong foundation in statistics, including hypothesis testing, confidence intervals, and experimental design Knowledge of reinforcement learning, planning, or decision-making systems Experience evaluating multi-modal or multi-agent systems Prior work on AI reliability, safety, or benchmarking

    About Company

    We bring amazing people together to make amazing things happen.

    We’re a diverse collection of thinkers and doers, continually reimagining what’s possible to help us all do what we love in new ways. The people who work here have reinvented entire industries with the Mac, iPhone, iPad, and Apple Watch, as well as with services, including iTunes, the App Store, Apple Music, and Apple Pay. And the same passion for innovation that goes into our products also applies to our practices — strengthening our commitment to leave the world better than we found it.

    About Apple

    There’s a place here for every kind of brilliant. Everyone here is an innovator, or an innovator-to-be, no matter what your team or your role. So bring your passion, courage, and original thinking and get ready to share it, because every new product, service, or feature we invent is the result of people working together to make each others’ ideas stronger. Innovation at this level depends on people who represent the variety of the human experience and inspire us with their own fresh perspectives. Together, we’ll do amazing work that can make a difference in people’s lives. Including your own. Learn more about working at Apple.

    Similar Jobs

    See more jobs