Jobs · Engineering

Sr. Principal Software Engineer

Cerence AI · United States · 2 wk ago
RemoteRemoteEngineering$185k/yrFull-time

About the Role

We're seeking an exceptional Senior Principal Software Engineer to join our global team and drive the future of mobility. As a key member of our team, you'll optimize and deploy high-performance LLM inference pipelines, own inference runtimes across various platforms, and push model performance through advanced techniques.

Responsibilities

  • Optimize and deploy high-performance LLM inference pipelines
  • Own inference runtimes across data center, edge, and embedded platforms
  • Push model performance through quantization, kernel fusion, and cache optimization
  • Drive latency and throughput improvements for production products
  • Enable efficient, reliable deployment without external vendor dependency
  • Build deep expertise in inference engines (vLLM, TensorRT-LLM, llama.cpp, QAIRT)
  • Implement and evaluate quantization strategies (INT8, INT4, FP4, FP8, mixed precision)
  • Optimize key-value cache performance through paging, prefix caching, and cache-aware memory layout design
  • Design and tune batching strategies, continuous batching, and speculative decoding for latency and throughput optimization

Required Experience & Skills

  • Proven experience optimizing ML inference performance in production
  • Deep understanding of GPU architecture and memory hierarchies
  • Hands-on experience with CUDA and low-level performance tuning
  • Experience deploying models beyond research environments
  • Expertise in inference engines (vLLM, TensorRT-LLM, llama.cpp, QAIRT)
  • CUDA kernel development and profiling
  • Quantization techniques (INT8/INT4/FP4/FP8, AWQ, GPTQ)
  • KV cache optimization and memory layout design
  • Latency optimization (batching, speculative decoding, continuous batching)

What Success Looks Like

  • Efficient model deployment on edge and embedded devices
  • Significant improvement in tokens/sec compared to baseline implementations
  • Minimized and predictable end-to-end latency
  • Material reduction in inference cost per request
  • Elimination of dependency on partners for inference optimization

Pay

Salary range: $185,000.00 USD - $280,000.00 USD. The actual salary will be determined based on experience and other job-related factors. It is not typical for offers to be made at or near the top of the range.

Benefits

  • Annual bonus opportunity
  • Insurance coverage (medical, dental, vision, life, and disability)
  • Paid time off
  • Paid holidays
  • Company contribution to the RRSP (Registered Retirement Savings Plan)
  • Equity awards for certain positions and levels
  • Remote and/or hybrid work available depending on the position

All compensation and benefits are subject to the terms and conditions of the underlying plans or programs, as applicable, and may be amended, terminated, or replaced from time to time.

About Cerence

Cerence is the global leader in AI for transportation, specialized in building AI and voice-powered companions for cars, two-wheelers, and more. With over 500 million cars shipped with Cerence AI's technology, we partner with leading automakers, mobility providers, and technology companies to power intuitive, integrated experiences that create safer, more connected, and more enjoyable journeys for drivers and passengers alike.

Our team is dedicated to pushing the boundaries of AI innovation, working globally with headquarters in Burlington, Massachusetts, USA, and 16 other offices across Europe, Asia, and North America. We bring together diverse backgrounds and varied skill sets with the shared goal of advancing the next generation of transportation user experiences.

Similar jobs