Apple Inc logo

Senior Software Engineer

Apple Inc

  • Cupertino, CA
  • 5 days ago

    Highlights

    transformer architectures, RLHF/RLAIF, fine-tuning, reward modeling ) and frameworks such as PyTorch or Hugging Face, as applied to evaluation and reward signal design Familiarity with eval-driven development - defining success criteria and test cases from product goals and real user workflows rather than abstract benchmarks Experience with data science methods applied to quality measurement - defining ground truth, measuring inter-rater agreement (e.g. correlation, confidence intervals, hypothesis testing) Experience with MLOps, deployment, and test/eval environment management - containerization, CI/CD, model versioning, monitoring, cloud platforms (AWS, GCP, or similar), and staging or provisioning environments to produce repeatable, deterministic conditions Comfort communicating and collaborating effectively across multicultural teams and time zones.

    Numbers & Facts

    LocationCupertino, CA
    IndustryComputer/IT Services
    Company Size10,000 employees or more
    Year Founded1976
    Websitehttps://www.apple.com/jobs

    Description

    As part of the Siri organization, you will build the systems and tooling that make evaluation a first-class part of how Siri is developed - not an after-the-fact check - spanning human evaluation, real user feedback, reward and alignment signals, and data science rigor across iOS, iPadOS, macOS, watchOS, and visionOS.

    This is a rare opportunity to work at the intersection of software engineering and rigorous evaluation science - applying machine learning engineering techniques, from model evaluation to reward modeling, to build the infrastructure that keeps this quality signal trustworthy. What you build will directly shape the direction of one of the worlds most widely used assistants.

    In this role youll contribute across several interconnected work streams spanning evaluation quality, reward/alignment signals, and data science. A core part of the job is bringing "evals first" thinking to the team - building tooling and harnesses grounded in real workflows, and designing for observability and reproducibility so quality can be measured clearly and issues caught early. This is a largely unexplored space with few established playbooks, so being self-driven is a must - youll define your own path as much as execute one. Scope and priorities will evolve, and were looking for someone who moves fluidly across these areas, bringing strong software engineering fundamentals with enough ML/LLM depth to build and ship AI-facing toolingSupporting the evaluation of new Siri features and interaction modalities, working from ambiguous early requirements toward concrete, automated coverage Turning product goals into measurable system behavior - instrumenting the product, building eval harnesses, and creating test datasets grounded in real user workflows Building and improving reward models and alignment signals that measure whether Siri responses meet user needs Designing and shipping evaluation tooling, pipelines, and architecture end-to-end - from data ingestion through scoring to monitoring in production - with an eye toward observability, logging, and reproducibility Diagnosing failures across the stack, from environment provisioning through pipeline execution to scoring - enabling auto-diagnostics and driving durable fixes by partnering across engineering, infrastructure, and program teams to align on interfaces, priorities, and shared standardsStrong programming skills in one or more compiled languages (Swift, C++, or Objective-C) Strong Python skills and solid computer science fundamentals, including data structures, algorithms, and clean, testable code Ability to quickly learn and adapt to evolving technologies and tools, such as GenAI-assisted coding, new ML frameworks, and emerging LLM/agent tooling Experience with backend/API development and production debugging Excellent communication and cross-team collaboration skills, with experience working effectively within large, cross-functional organizations M.S. or B.S. in Computer Science, Machine Learning, or a related field (or equivalent experience)Experience evaluating ML, LLM, or agent-based systems, including familiarity with metrics, scoring methodology, trajectory and outcome analysis, and techniques like prompting, RAG, or LLM as judge Understanding of reinforcement learning and the underlying techniques behind modern LLMs (e.g. transformer architectures, RLHF/RLAIF, fine-tuning, reward modeling ) and frameworks such as PyTorch or Hugging Face, as applied to evaluation and reward signal design Familiarity with eval-driven development - defining success criteria and test cases from product goals and real user workflows rather than abstract benchmarks Experience with data science methods applied to quality measurement - defining ground truth, measuring inter-rater agreement (e.g. Cohens/Fleiss kappa), and validating automated scorers using basic statistical techniques (e.g. correlation, confidence intervals, hypothesis testing) Experience with MLOps, deployment, and test/eval environment management - containerization, CI/CD, model versioning, monitoring, cloud platforms (AWS, GCP, or similar), and staging or provisioning environments to produce repeatable, deterministic conditions Comfort communicating and collaborating effectively across multicultural teams and time zones

    About Company

    We bring amazing people together to make amazing things happen.

    We’re a diverse collection of thinkers and doers, continually reimagining what’s possible to help us all do what we love in new ways. The people who work here have reinvented entire industries with the Mac, iPhone, iPad, and Apple Watch, as well as with services, including iTunes, the App Store, Apple Music, and Apple Pay. And the same passion for innovation that goes into our products also applies to our practices — strengthening our commitment to leave the world better than we found it.

    About Apple

    There’s a place here for every kind of brilliant. Everyone here is an innovator, or an innovator-to-be, no matter what your team or your role. So bring your passion, courage, and original thinking and get ready to share it, because every new product, service, or feature we invent is the result of people working together to make each others’ ideas stronger. Innovation at this level depends on people who represent the variety of the human experience and inspire us with their own fresh perspectives. Together, we’ll do amazing work that can make a difference in people’s lives. Including your own. Learn more about working at Apple.

    Similar Jobs

    See more jobs