Staff Electrical Quality Engineer Symbotic
- $125,000–$171,600 Per Year
| Location | Cambridge, MA |
| Industry | Computer/IT Services |
| Company Size | 10,000 employees or more |
| Year Founded | 1976 |
| Website | https://www.apple.com/jobs |
Join the team redefining what a deeply personal and integrated assistant can be.
As part of the Siri organization, you will help shape one of the worlds most widely used AI assistants, powered by our next-generation of Apple Intelligence, with capabilities like personal context understanding and on-screen awareness, built with privacy from the ground up. Your work will have direct, meaningful impact for users across iOS, iPadOS, macOS, watchOS, and visionOS.
Our Speech Evaluation team sits at the center of Apples ASR, TTS, and real-time conversational AI efforts, partnering directly with the modeling teams. Were growing the team to take on a role focused specifically on evaluating audio LLMs: designing the datasets that stress-test them and the metrics that decide whether theyre ready. Youll help define how Apple measures a new class of models that listen, speak, and reason.
This is a rare opportunity to build at the intersection of cutting-edge AI and human-centered design, shipping technology that is centered around users and their needs. This role owns the data and metrics foundation for evaluating speech LLMs (e.g., real-time speech understanding and generation models) across accuracy, robustness, and conversational quality. Youll build and curate evaluation datasets that reflect real usage - from personalized named-entity queries to multi-turn fluid conversations - and design the metrics and automated judges that turn model outputs into actionable, trustworthy signal. Youll work closely with modeling, infrastructure, and product partners to make sure every new model is evaluated quickly, consistently, and at the right level of rigor before it reaches customers. Designs and curates audio evaluation datasets that represent real-world usage, including personalized, multilingual, and conversational scenarios. Defines and implements evaluation metrics for audio LLMs, spanning accuracy, robustness, and conversational/generation quality. Builds automated evaluation pipelines and LLM-as-judge tooling to scale audio model assessment without sacrificing reliability. Analyzes model evaluation results to identify accuracy gaps, regressions, and opportunities for hillclimbing, and communicates findings to modeling teams. Partners with human-evaluation programs to design rating protocols and validate that automated metrics correlate with human judgment. Collaborates with infrastructure teams to integrate new evaluation sets and metrics into shared tooling. Contributes evaluation methodology for new audio LLM capabilities as they emerge, adapting existing frameworks to novel model behaviors.Bachelors degree in Computer Science, Electrical Engineering, or a related field, or equivalent practical experience. Experience building or working with text, speech or audio evaluation pipelines and metrics. Proficiency in Python and experience building data processing pipelines at scale. Experience curating or annotating datasets for machine learning evaluation or training. Working knowledge of statistics as applied to measuring model performance and interpreting evaluation results. Familiarity with large language model evaluation techniques, including automated (LLM-as-judge) and human evaluation methods. Strong written and verbal communication skills, with the ability to explain evaluation results to both technical and non-technical audiences.Experience evaluating audio-native or multimodal (speech-in, speech-out) large language models. Experience designing or running human evaluation studies (e.g., side-by-side comparisons, MOS ratings) at scale. Familiarity with personalization and named-entity evaluation challenges in speech systems. Experience with multilingual or international audio dataset development. Experience with distributed data processing frameworks (e.g., Spark) for large-scale audio dataset generation. Publication record or demonstrated contributions in speech, audio ML, or NLP evaluation.
We’re a diverse collection of thinkers and doers, continually reimagining what’s possible to help us all do what we love in new ways. The people who work here have reinvented entire industries with the Mac, iPhone, iPad, and Apple Watch, as well as with services, including iTunes, the App Store, Apple Music, and Apple Pay. And the same passion for innovation that goes into our products also applies to our practices — strengthening our commitment to leave the world better than we found it.
There’s a place here for every kind of brilliant. Everyone here is an innovator, or an innovator-to-be, no matter what your team or your role. So bring your passion, courage, and original thinking and get ready to share it, because every new product, service, or feature we invent is the result of people working together to make each others’ ideas stronger. Innovation at this level depends on people who represent the variety of the human experience and inspire us with their own fresh perspectives. Together, we’ll do amazing work that can make a difference in people’s lives. Including your own. Learn more about working at Apple.