Jobs · Engineering · California

Machine Learning Engineer, Speech LLM Training

Jobright.ai · San Francisco, CA · 1 wk ago
HybridEngineering$200k–$540k/yrFull-time

This role is part of the Jobright TNT—a private hiring network connecting top talent with leading AI startups such as Perplexity, Mercor, Cresta, Suno, and 150+ others. Only select, high-signal candidates are invited and recommended directly to hiring teams.

About the Company

Plaud Inc. is a hardware and software company building AI-powered voice recorders for note-taking, with over 1 million users globally.

Why Join Us

Compensation range: $200K/yr - $540K/yr.

Responsibilities

  • Build and train large-scale audio or speech models from the ground up, including unified SpeechLLMs, advanced ASR, expressive TTS, or generative audio architectures.
  • Work at the intersection of research and engineering, designing novel sequence modeling architectures and debugging distributed training clusters.
  • Traverse the entire stack—from fundamental signal processing and raw acoustic representations to massive foundation model training and edge-device optimization.
  • Demonstrate deep expertise in PyTorch or JAX, optimizing large-scale distributed training runs, managing GPU memory utilization, and resolving complex performance bottlenecks.
  • Thrive in a fast-paced, high-growth startup environment, taking extreme ownership of ambiguous problems and driving them into production.
  • Build AI systems that natively understand and generate speech, creating a hardware-software AI companion to amplify human productivity.

Requirements

  • Proven track record of building and training large-scale audio or speech models from the ground up, including unified SpeechLLMs, advanced ASR, expressive TTS, or generative audio architectures.
  • Experience working at the intersection of research and engineering, designing novel sequence modeling architectures and debugging distributed training clusters.
  • Comfort traversing the entire stack—from fundamental signal processing and raw acoustic representations to massive foundation model training and edge-device optimization.
  • Deep expertise in PyTorch or JAX, with experience optimizing large-scale distributed training runs, managing GPU memory utilization, and resolving complex performance bottlenecks.
  • Ability to thrive in a fast-paced, high-growth startup environment, taking extreme ownership of ambiguous problems and driving them into production.
  • Passion for building AI systems that natively understand and generate speech, ultimately creating a hardware-software AI companion.

Preferred Qualifications

  • Text-based LLMs: Hands-on experience with core text-based Large Language Model pretraining, instruction tuning, or RLHF.
  • Neural Audio Codecs: Hands-on experience designing and training state-of-the-art neural audio codecs for streamable, high-fidelity audio.
  • Generative Architectures: Designing and training diffusion models, flow matching, or autoregressive architectures specifically for speech and voice generation.
  • Alignment & Steerability: Applying Reinforcement Learning (RL) techniques (e.g., RLHF or GRPO) to improve conversational cadence, steerability, and alignment in foundation models.
  • Deep System Optimization: End-to-end inference and performance optimization using high-throughput serving frameworks (e.g., vLLM, TensorRT-LLM, SGLang) to minimize latency for real-time cloud streaming.
  • Large-Scale Infrastructure: Managing massive GPU clusters, utilizing advanced distributed training frameworks (e.g., FSDP, DeepSpeed), and navigating orchestration tools like Kubernetes.

Similar jobs