Jobs · Engineering · California

Machine Learning Engineer, Inference & Serving (Speech LLM)

Jobright.ai · San Francisco, CA · 1 wk ago
HybridEngineering$195k–$365k/yrFull-time

This role is part of Jobright TNT—the private hiring network connecting top talent with top AI startups like Perplexity, Mercor, Cresta, Suno, and 150+ more. Only select, high-signal candidates are invited to Jobright TNT and recommended directly to hiring teams.

About the Company

Plaud is a hardware and software company making AI-powered voice recorders for note-taking. The best-selling AI note-taking device globally, with 700K units shipped across 170+ countries. Bootstrapped to $180M ARR with 10x YoY growth for two consecutive years, no outside funding. 400+ person team led by serial entrepreneurs, with offices in San Francisco, Tokyo, and Singapore.

Pay

$195K/yr - $365K/yr

Why Join Us

  • Competitive compensation and benefits, with flexibility for strong candidates

Responsibilities

  • Build and deploy high-throughput, ultra-low-latency inference engines for large language models or foundational speech models
  • Navigate tradeoffs between latency, throughput, and Time-To-First-Token (or Time-To-First-Audio) in real-time streaming environments
  • Implement continuous batching, KV cache management (e.g., PagedAttention), and stateful connections for real-time conversational AI
  • Optimize GPU architectures (NVIDIA Ampere/Hopper) and memory hierarchy to eliminate hardware bottlenecks
  • Collaborate at the intersection of the core ML training team and backend infrastructure team
  • Thrive in fast-moving environments, focusing on systems-engineering challenges to maximize GPU cluster performance
  • Develop AI systems that natively understand and generate speech, creating a hardware-software AI companion for human productivity

Requirements

  • Hands-on experience building and deploying high-throughput, ultra-low-latency inference engines for LLMs or foundational speech models
  • Deep understanding of latency, throughput, and Time-To-First-Token (or Time-To-First-Audio) tradeoffs in real-time streaming
  • Practical experience with continuous batching, KV cache management (e.g., PagedAttention), and stateful connections for conversational AI
  • Expertise in GPU architectures (NVIDIA Ampere/Hopper) and memory hierarchy optimization
  • Strong communication and collaboration skills
  • Ability to thrive in fast-moving environments
  • Passion for building AI systems that understand and generate speech

Preferred Qualifications

  • Frontier Serving Frameworks: Deep familiarity with modern LLM serving frameworks like vLLM, TensorRT-LLM, SGLang, or NVIDIA Triton Inference Server (bonus for open-source contributions)
  • Real-Time Audio Streaming: Experience handling continuous audio streams over WebSockets or WebRTC, deploying neural audio codecs, and managing chunked audio generation to minimize latency
  • Advanced Inference Techniques: Implementation of speculative decoding, lookahead decoding, or chunked prefill
  • Model Compression & Quantization: Hands-on experience with post-training quantization (PTQ), FP8, INT8, AWQ, or GPTQ without degrading audio naturalness or ASR accuracy
  • Large-Scale Distributed Systems: Experience deploying multi-GPU (Tensor Parallelism) and multi-node inference pipelines, and managing autoscaling infrastructure using Kubernetes

Similar jobs