Machine Learning Engineer, Inference & Serving (Speech LLM)
Jobright.ai · San Francisco, CA · 1 wk ago
HybridEngineering$195k–$365k/yrFull-time
This role is part of Jobright TNT—the private hiring network connecting top talent with top AI startups like Perplexity, Mercor, Cresta, Suno, and 150+ more. Only select, high-signal candidates are invited to Jobright TNT and recommended directly to hiring teams.
About the Company
Plaud is a hardware and software company making AI-powered voice recorders for note-taking. The best-selling AI note-taking device globally, with 700K units shipped across 170+ countries. Bootstrapped to $180M ARR with 10x YoY growth for two consecutive years, no outside funding. 400+ person team led by serial entrepreneurs, with offices in San Francisco, Tokyo, and Singapore.
Pay
$195K/yr - $365K/yr
Why Join Us
- Competitive compensation and benefits, with flexibility for strong candidates
Responsibilities
- Build and deploy high-throughput, ultra-low-latency inference engines for large language models or foundational speech models
- Navigate tradeoffs between latency, throughput, and Time-To-First-Token (or Time-To-First-Audio) in real-time streaming environments
- Implement continuous batching, KV cache management (e.g., PagedAttention), and stateful connections for real-time conversational AI
- Optimize GPU architectures (NVIDIA Ampere/Hopper) and memory hierarchy to eliminate hardware bottlenecks
- Collaborate at the intersection of the core ML training team and backend infrastructure team
- Thrive in fast-moving environments, focusing on systems-engineering challenges to maximize GPU cluster performance
- Develop AI systems that natively understand and generate speech, creating a hardware-software AI companion for human productivity
Requirements
- Hands-on experience building and deploying high-throughput, ultra-low-latency inference engines for LLMs or foundational speech models
- Deep understanding of latency, throughput, and Time-To-First-Token (or Time-To-First-Audio) tradeoffs in real-time streaming
- Practical experience with continuous batching, KV cache management (e.g., PagedAttention), and stateful connections for conversational AI
- Expertise in GPU architectures (NVIDIA Ampere/Hopper) and memory hierarchy optimization
- Strong communication and collaboration skills
- Ability to thrive in fast-moving environments
- Passion for building AI systems that understand and generate speech
Preferred Qualifications
- Frontier Serving Frameworks: Deep familiarity with modern LLM serving frameworks like vLLM, TensorRT-LLM, SGLang, or NVIDIA Triton Inference Server (bonus for open-source contributions)
- Real-Time Audio Streaming: Experience handling continuous audio streams over WebSockets or WebRTC, deploying neural audio codecs, and managing chunked audio generation to minimize latency
- Advanced Inference Techniques: Implementation of speculative decoding, lookahead decoding, or chunked prefill
- Model Compression & Quantization: Hands-on experience with post-training quantization (PTQ), FP8, INT8, AWQ, or GPTQ without degrading audio naturalness or ASR accuracy
- Large-Scale Distributed Systems: Experience deploying multi-GPU (Tensor Parallelism) and multi-node inference pipelines, and managing autoscaling infrastructure using Kubernetes