About the Role
We're looking for a Member of Technical Staff, Voice focused on voice and audio to help advance the spoken intelligence behind Pi. In this role, you'll work at the intersection of research and production-developing, training, and shipping neural models across the full spectrum of voice: speech synthesis, recognition, audio generation, and real-time spoken dialogue. You'll collaborate closely with ML engineers, product teams, and infrastructure to turn cutting-edge ideas in areas like neural audio codecs, diffusion-based TTS, and multimodal foundation models into the natural, expressive voice experiences that millions of Pi users interact with every day.
What You'll Do
- Research, develop, and optimize neural models for voice and audio-including text-to-speech, automatic speech recognition, audio generation, and spoken dialogue systems.
- Build and maintain production-grade training and inference pipelines for voice models, with close attention to latency, naturalness, and scalability.
- Run experiments end-to-end: data curation, model architecture design, training, evaluation, and ablation studies.
- Collaborate with ML engineers, product teams, and infrastructure to integrate voice models into Pi's real-time conversational stack.
- Explore and apply advances in neural audio codecs, diffusion-based synthesis, streaming architectures, and multimodal foundation models to improve Pi's voice experience.
- Develop robust evaluation frameworks combining perceptual metrics, automated benchmarks, and user-facing quality signals.
- Contribute to Inflection's research culture through publications, internal reviews, and knowledge sharing.
What We're Looking For
- 2-5 years of research or engineering experience (including graduate work) in audio, speech, or multimodal ML.
- Strong proficiency in PyTorch and hands-on experience training and debugging large-scale neural models on GPU/accelerator clusters.
- Solid understanding of audio and speech fundamentals spectrograms, mel features, vocoders, codec-based representations, and signal processing.
- Demonstrated ability to take a research idea from prototype to production: equally comfortable reading papers and writing efficient, CUDA-aware training loops.
- Familiarity with modern generative architectures for audio (e.g., diffusion models, autoregressive codecs, flow-matching) and their trade-offs.
- Clear, collaborative communication able to distill complex research into actionable insights for cross-functional partners.
- Have a bachelor's degree or equivalent in Computer Science, Electrical Engineering, Linguistics, or a related field; MS or PhD strongly preferred.
Employee Pay Disclosures
At Inflection AI, we aim to attract and retain the best employees and compensate them in a way that appropriately and fairly values their individual contributions to the company. For this role, Inflection AI estimates a starting annual base salary to fall within the range of $225,000 to $325,000, depending on a candidate's qualifications and level of experience. This role also includes a meaningful equity component, allowing employees to share in the long-term success of the company.
Benefits
Inflection AI values and supports our team's mental, emotional, financial and physical health. We are focused on building a positive, safe, inclusive and inspiring place to work. Our benefits include:
- Robust medical, dental and vision options with employer contributions for HSA, FSA and DFSA
- 401k matching program
- Flexible Time Off, 10 paid holidays, 5 days sick leave
- Parental, Medical and Family care leave
- Generous cell-phone, wellness and office set up stipends
- Support of country-specific visa needs for international employees living in the Bay Area