Jobs · Engineering · California

AI Inference Engineer - Speech

Maven · San Jose, CA · 1 wk ago
Engineering$152k/yrFull-time

Responsibilities

  • Developing state-of-the-art speech services for Zoom products.
  • Devising novel techniques where off-the-shelf solutions are not available.
  • Optimizing ASR inference systems for production deployment, including inference latency, throughput, memory footprint, and resource utilization.
  • Optimizing model inference performance by diving deep into the lower stack of inference frameworks, with a focus on hardware-specific optimizations for Nvidia GPUs.
  • Proposing new model structures by joint optimization of model accuracy and inference speed.
  • Designing and developing ASR systems with low latency and high accuracy requirements, while ensuring scalability of GPU infrastructure and improving throughput of ASR service.
  • Profiling and debugging ASR runtime performance bottlenecks across different deployment hardware and environments.

Requirements

  • Possess a Master's in Computer Science, Electrical Engineering or related fields with 3+ years of experience in speech recognition, speech-llm or AI model inference.
  • Show knowledge in deep learning and hands-on programming skills in Python, shell scripts, C/C++; familiarity with ML frameworks such as PyTorch and TensorFlow.
  • Demonstrate deep understanding of transformer encoder-decoder frameworks for speech recognition, including attention mechanisms, beam search and sequence-to-sequence modeling for end-to-end ASR systems.
  • Understand recent advancements in speech foundation models and speech-LLMs that integrate acoustic and linguistic representations, enabling unified modeling for speech understanding and transcription tasks.
  • Have experience in optimizing deep learning model inference on NVIDIA GPUs, including profiling and accelerating AI models using CUDA, TensorRT, and mixed-precision computation to achieve low latency, high-throughput performance.
  • Have experience developing and tuning custom CUDA kernels, leveraging CUDA Graphs for efficient execution scheduling, and minimizing kernel launch overhead to maximize GPU utilization.
  • Be proficient in end-to-end performance analysis, memory optimization, and deployment of largescale ML models on GPU clusters.
  • Experienced with stream management, asynchronous execution, and integrating frameworks such as PyTorch and TensorFlow for real-time inference.

Qualifications

  • Minimum Salary Range or On Target Earnings: $151,800.00 Maximum $332,200.00
  • Note: Starting pay will be based on a number of factors and commensurate with qualifications & experience.
  • We also have a location based compensation structure; there may be a different range for candidates in this and other locations.

Skills

  • Deep Learning
  • Speech Recognition
  • Model Inference
  • Transformer Encoder-Decoder Frameworks
  • NVIDIA GPUs
  • CUDA
  • TensorRT
  • Mixed-Precision Computation
  • Stream Management
  • Asynchronous Execution
  • PyTorch
  • TensorFlow

Benefits

  • Supports work-life balance
  • Community contribution opportunities
  • Flexible work arrangements
  • Health and wellness programs

Pay

USD 151,800-332,200 / year + Equity

Schedule

Not specified

Similar jobs

Voice AI Engineer

CleraSan Francisco, CA· 1 wk ago
Engineering$72k/yrapply on jobs.ashbyhq.com