Lead Product Manager
About the role
Own product direction across the full model lifecycle — data, training and adaptation, evaluation, release, production monitoring, and improvement or retirement — for our SLMs, ASR stack, and the real-time inference infrastructure that serves them. Retirement is a real part of that: the leading labs deliberately sunset models to concentrate effort, and we'd rather run a few models well than maintain a legacy model zoo.
Own the data strategy underneath it all: acquisition, consent and usage rights, sampling, and annotation. Model quality is decided here before the first training run — get the data model right and every ASR and SLM effort downstream gets simpler and better. On a platform built on customer conversations, consent and rights are foundational, not paperwork.
Turn ambiguous model-quality questions into decisions. "Transcripts got worse this week" is a starting point, not a ticket. You'll define what good means, get it measured, and decide what ships.
Sit inside eval reviews, error analyses, and incident retros as a peer. You should be able to look at a failing conversation trace and form your own hypothesis before the team tells you theirs.
Treat internal teams as customers. The voice agents, agentic runtime, and AI features across the product all run on your stack — product engineering needs model capabilities and latency/cost envelopes they can plan around, and GTM needs a roadmap you won't have to walk back.
Make trade-off calls with real constraints: model quality vs. streaming latency, train vs. fine-tune vs. buy, model size vs. capability, GPU cost vs. what the price point can absorb. These are the daily currency of this role, not edge cases.
Write. Direction memos, decision docs, and specs that engineers actually read. If your best work happens in slide decks, this isn't the right fit.
Release and rollback calls: whether a model ships, against quality bars you define.
The model roadmap and its sequencing — including what gets deprecated and when.
Where data investment goes: acquisition, annotation, and labeling priorities.
The quality bar itself: what "good enough" means for an ASR or SLM release, and how it's measured.
Skills
- A hands-on track record with models themselves. You've built or run models in production — trained, fine-tuned, served, or optimized them — not just orchestrated APIs around them. Building beats running. Our core work is custom SLMs, ASR, and the inference infrastructure behind them, so experience with speech or with models under real-time constraints counts double.
- A CS/ML degree alone doesn't count; neither does "worked closely with data scientists."
- Two years of product ownership, formally titled or not. You've been accountable for what got built and whether it worked, not just for the backlog.
- Fluency across the model and serving stack. You have informed opinions on eval design, when to fine-tune vs. train vs. distill, quantization and serving trade-offs, why WER alone is a lousy ASR metric, and what actually drives real-time inference cost. Opinions you can defend to someone who does this full-time.
- Judgment under uncertainty. Model behavior is probabilistic; roadmaps aren't. You can commit to outcomes without pretending the uncertainty away.
- Direct communication. You say what you think, change your mind when the evidence says so, and put decisions in writing.
Nice to have
- Speech experience specifically: training or productionizing ASR/TTS, telephony, streaming latency work.
- You've built training data pipelines or run labeling operations — sourcing, sampling, annotation quality, data rights.
- You've run inference infrastructure at scale — GPU capacity planning, serving optimization, cost-per-call tuning.
- You've built or run an eval harness in production, not just read about them.
- Experience pricing or packaging AI products.
- Publications, open-source work, or a technical blog we can read before we talk.