Jobs · Information Technology · California

Member of Technical Staff - Multi-Modal, Audio

Liquid AI · San Francisco, CA · 3 wk ago
HybridInformation TechnologyFull-time

About the role

Spun out of MIT CSAIL, Liquid AI builds general-purpose AI systems that run efficiently across deployment targets—from data center accelerators to on-device hardware—ensuring low latency, minimal memory usage, privacy, and reliability. The Audio team is building frontier speech-language models that handle STT, TTS, and speech-to-speech in a single architecture. This role sits at the center of applied audio model development, working directly with the technical lead to ship production systems that run on-device under real-time constraints. You will own critical workstreams across data pipelines, evaluation systems, and customer deployments.

Responsibilities

  • Build and scale data pipelines for audio model training, including preprocessing, augmentation, and quality filtering at scale
  • Design, implement, and maintain evaluation systems that measure multimodal performance across internal and public benchmarks
  • Fine-tune and adapt audio models for customer-specific use cases, owning delivery from requirements through deployment
  • Contribute production code to the core audio repository, collaborating with infrastructure and research teams
  • Support experimentation under real hardware constraints, shifting between customer work and core development as priorities evolve

Requirements

Must-have

  • Strong programming fundamentals with demonstrated ability to write clean, maintainable, production-grade code
  • Experience building and shipping production ML systems beyond model training (data pipelines, evals, serving infrastructure)
  • Proficiency in PyTorch and familiarity with distributed training frameworks (DeepSpeed, FSDP, or similar)
  • Track record of collaborating effectively in shared codebases with high engineering standards

Nice-to-have

  • Direct experience with audio/speech models (ASR, TTS, vocoders, diarization, or speech-to-speech systems)
  • Experience designing and running large-scale training experiments on distributed GPU clusters
  • Open-source contributions that demonstrate code quality and engineering judgment

What success looks like

  • Within 6 months, you independently deliver production-ready data pipelines or evaluation systems and own at least one customer workstream end-to-end
  • Your PRs to the core audio repo are accepted without heavy rework, demonstrating strong judgment in system design
  • By year end, you operate as a second pillar to the technical lead, unblocking parallel workstreams and raising overall team velocity

Benefits

  • Rare technical problems: Work on audio-to-audio frontier systems with real ownership in a team small enough that your contributions ship directly to production
  • Competitive base salary with equity in a unicorn-stage company
  • We pay 100% of medical, dental, and vision premiums for employees and dependents
  • 401(k) matching up to 4% of base pay
  • Unlimited PTO plus company-wide Refill Days throughout the year

Similar jobs