Job Description
Job Description: Principal Research Scientist Speech & Audio Foundation Models
Text-to-speech, voice cloning, speech synthesis, realtime conversational voice.
$270,000-$500,000 base plus bonus, equity and benefits (US).
Relocation assistance available. Visa transfer supported
San Francisco on-site preferred Remote considered in the US, UK and parts of Europe.
Permanent, full-time.
A top end research lab building realtime voice models text-to-speech, speech-to-text and speech-to-speech delivered as an API. The models run in production behind consumer applications used at very large scale, across health, learning, therapy, companionship, media and gaming.
Text-to-speech, voice cloning, speech synthesis, realtime conversational voice.
The role
Build the models that are the product! This is full-stack research ownership: you frame the question, run the experiments, and ship the result. Research is only finished when it is in production and measurable.
Responsibilities
Train foundation models: pre-training, reinforcement learning, reward modelling, post-training, new architectures, scaling.
Build and improve voice and speech models across TTS, STT and speech-to-speech.
Design the evaluation that proves the work: benchmarks, eval loops, LLM-as-judge, failure analysis. Evaluation is treated as a research product in its own right, not as a pre-launch checkbox.
Work on frontier problems adjacent to the roadmap: multimodal, agents and tool use, test-time compute.
Take models into production alongside the serving engineering team, inside a sub-200ms latency budget and across 100+ languages.
Essential
Hands-on foundation-model training. Pre-training, RL, reward modelling, post-training, scaling. Fine-tuning or building on top of someone else's model is a different discipline and is not what this role is.
Real voice or speech research: TTS, STT or speech-to-speech. Speech-to-speech is the strongest signal; TTS and ASR both count. Text-only research does not transfer.
Evidence you can point at: papers, shipped models, open-source contributions, or systems in production.
Desirable
Evaluation depth: benchmarks, eval loops, quality measurement, failure analysis.
Publications at ICML, ICLR, NeurIPS, EMNLP, ACL, AAAI, Interspeech or ICASSP.
PhD in ML or NLP, or equivalent practical experience you can point to.
Frontier exposure: multimodal, agents, tool use, test-time compute.
Public work: side projects, open-source, technical write-ups.