Staff Research Scientist
About the role
Build expressive speech foundation models at trillion-parameter scale, with real world impact. This team has built their entire speech stack in-house, including proprietary LLM-based ASR and TTS, both already outperforming SOTA benchmarks. They're already powering real-time, human-to-AI conversations at scale, every day. Now they're working at trillion-parameter scale to push toward a genuine end-to-end speech-to-speech LLM, one that understands and responds with genuine emotional intelligence and able to have natural, human-like conversations. That means solving problems like long-context reasoning, pronunciation accuracy and maintaining consistency in noisy, real-world environments, in a domain where no model has cracked this yet.
Responsibilities
- Build SOTA speech models from the ground up, at genuinely large scale
- Own problems end-to-end, from research through to production
- Solve hard, domain-specific speech challenges as part of the push toward speech-to-speech LLM research
- Help shape technical direction as part of a team working at the frontier of this space
Requirements
- Deep, hands-on expertise in at least one of: Speech/Audio LLMs, TTS or audio generation, speech-to-speech or large-scale speech understanding
- Experience shipping speech systems at real scale, not just research-stage prototypes
Nice to have
- Experience with pre-training speech foundation models (HuBERT, Wav2Vec or similar)
- Experience with Multimodal LLMs especially to reduce hallucinations
- Experience with Speech LLMs for low latency streaming ASR
Pay
Up to $350K–$400K base (DOE), plus substantial equity.
Schedule
This is an onsite position in South Bay or Seattle.