Research Scientist-Voice and Audio Ai

RST Recruitment

  • Redwood City, California
  • 24 days ago

    Highlights

    Strong ML Research Background – You've worked on advanced ML problems (for example: LLM pre-training and post training, transcription model training, text to speech model training, or multimodal systems), either in industry or academia. Research & Experimentation – Explore and develop new techniques across LLMs and audio models to improve reasoning, latency, and conversational quality in real-time systems.

    Numbers & Facts

    LocationRedwood City, California

    Description

    This is a research-driven, high-impact role for ML researchers who want to push the boundaries of real-time AI. As a Founding Machine Learning Research Engineer at Retell, you'll focus on advancing model capabilities for human-like voice agents operating in complex, real-world environments.


    You'll explore new approaches across LLMs and audio models, design novel evaluation methods, and prototype systems that improve reasoning, latency, and conversational quality. Your work will directly influence production systems, bridging cutting-edge research with real-world deployment.


    If you're excited about solving open-ended ML problems, experimenting rapidly, and shaping how voice AI systems think and perform, this is a unique opportunity to do so at scale.


    KEY RESPONSIBILITIES

    • Research & Experimentation – Explore and develop new techniques across LLMs and audio models to improve reasoning, latency, and conversational quality in real-time systems.
    • Model Prototyping – Rapidly build and iterate on experimental models and pipelines, turning research ideas into working prototypes.
    • Evaluation & Benchmarking – Design novel evaluation frameworks, datasets, and metrics to measure performance on complex, real-world voice tasks.
    • Bridge Research to Production – Collaborate closely with engineering to translate research insights into deployable systems.
    • Human Feedback Loops – Develop methods to incorporate human evaluation into model improvement, especially for subjective conversational quality.
    • Advance the Frontier – Stay at the cutting edge of ML research and bring new ideas into Retell's product and infrastructure.

    HOW TO THRIVE

    • Strong ML Research Background – You've worked on advanced ML problems (for example: LLM pre-training and post training, transcription model training, text to speech model training, or multimodal systems), either in industry or academia.
    • Deep Technical Foundation – Comfortable with PyTorch, model architectures, and the math behind modern machine learning.
    • Experimental Mindset – You enjoy exploring open-ended problems and iterating quickly on ideas.
    • Bridging Theory & Practice – You can translate research into systems that work in real-world environments.
    • Startup-Ready – You thrive in fast-paced environments with high ownership and ambiguity.
    • Collaborative & Clear Communicator – You can explain complex ideas and work cross-functionally to drive impact.

    Tech stack:
    PyTorch, LLMs, Audio/Speech Models, Text-to-Speech (TTS), Automatic Speech Recognition (ASR), Multimodal Systems, Python

    Seniority:

    1 - 5 years of experience in audio/multimodal AI research or engineering

    Work experience:

    Working at a high bar company with AI products (MAANG, vc backed startup, etc.)

    Experience with LLM pre-training or post-training, evals, and translating research to production

    Coming from another top voice ai or audio startup (Cartesia, Eleven Labs, Descript, etc.)

    Education:

    Degree in CS, ML, or closely related field (PhD prefered)

    Recent publications in voice, audio, or speech AI

    Hard skills:

    Hands-on PyTorch and audio/speech model development

    Experience with TTS, ASR, or multimodal audio systems

    Pre-training experience at scale

    Miscellaneous:

    Comfortable with intense startup pace including weekend work

    Similar Jobs

    See more jobs