AI Inference Engineer - Speech
Maven · San Jose, CA · 1 wk ago
Engineering$152k/yrFull-time
Responsibilities
- Developing state-of-the-art speech services for Zoom products.
- Devising novel techniques where off-the-shelf solutions are not available.
- Optimizing ASR inference systems for production deployment, including inference latency, throughput, memory footprint, and resource utilization.
- Optimizing model inference performance by diving deep into the lower stack of inference frameworks, with a focus on hardware-specific optimizations for Nvidia GPUs.
- Proposing new model structures by joint optimization of model accuracy and inference speed.
- Designing and developing ASR systems with low latency and high accuracy requirements, while ensuring scalability of GPU infrastructure and improving throughput of ASR service.
- Profiling and debugging ASR runtime performance bottlenecks across different deployment hardware and environments.
Requirements
- Possess a Master's in Computer Science, Electrical Engineering or related fields with 3+ years of experience in speech recognition, speech-llm or AI model inference.
- Show knowledge in deep learning and hands-on programming skills in Python, shell scripts, C/C++; familiarity with ML frameworks such as PyTorch and TensorFlow.
- Demonstrate deep understanding of transformer encoder-decoder frameworks for speech recognition, including attention mechanisms, beam search and sequence-to-sequence modeling for end-to-end ASR systems.
- Understand recent advancements in speech foundation models and speech-LLMs that integrate acoustic and linguistic representations, enabling unified modeling for speech understanding and transcription tasks.
- Have experience in optimizing deep learning model inference on NVIDIA GPUs, including profiling and accelerating AI models using CUDA, TensorRT, and mixed-precision computation to achieve low latency, high-throughput performance.
- Have experience developing and tuning custom CUDA kernels, leveraging CUDA Graphs for efficient execution scheduling, and minimizing kernel launch overhead to maximize GPU utilization.
- Be proficient in end-to-end performance analysis, memory optimization, and deployment of largescale ML models on GPU clusters.
- Experienced with stream management, asynchronous execution, and integrating frameworks such as PyTorch and TensorFlow for real-time inference.
Qualifications
- Minimum Salary Range or On Target Earnings: $151,800.00 Maximum $332,200.00
- Note: Starting pay will be based on a number of factors and commensurate with qualifications & experience.
- We also have a location based compensation structure; there may be a different range for candidates in this and other locations.
Skills
- Deep Learning
- Speech Recognition
- Model Inference
- Transformer Encoder-Decoder Frameworks
- NVIDIA GPUs
- CUDA
- TensorRT
- Mixed-Precision Computation
- Stream Management
- Asynchronous Execution
- PyTorch
- TensorFlow
Benefits
- Supports work-life balance
- Community contribution opportunities
- Flexible work arrangements
- Health and wellness programs
Pay
USD 151,800-332,200 / year + Equity
Schedule
Not specified