Jobs · Engineering

INFERENCE OPTIMIZATION ENGINEER

Up Top · United States · 1 wk ago
RemoteRemoteEngineeringFull-time

About the role

This is a hybrid IC / leadership role reporting to the Head of Engineering. You'll own the technical strategy for inference performance at scale, then build and lead a dedicated Inference Optimization team. The mandate is simple to state and hard to execute: drive down latency, increase throughput, and cut cost-per-token for large-scale LLM inference across a modern GPU fleet.

What you'll do

  • Own the end-to-end technical strategy for inference performance across the platform
  • Recruit, build, and lead the Inference Optimization team
  • Optimize GPU infrastructure across current and next-gen architectures (e.g. H200, B300)
  • Improve latency, throughput, and cost-per-token for production LLM inference workloads
  • Build reproducible benchmarking harnesses across inference engines such as vLLM and SGLang
  • Tune inference routing and multivariate load-balancing algorithms
  • Evaluate emerging optimization techniques — custom CUDA/Triton kernels, attention variants, quantization schemes, and compilation improvements
  • Assess emerging inference hardware (FPGAs, ASICs, and custom silicon)

What we're looking for

  • 8+ years in performance optimization or HPC, with deep GPU architecture and parallel-programming expertise
  • 5+ years leading engineering teams
  • Proficiency in Python, Rust, or Go
  • Hands-on experience running production LLM inference engines at high volume
  • Depth in modern inference optimization: continuous batching, PagedAttention / KV-cache management, speculative decoding, quantization, CUDA graphs, and torch.compile
  • A strong grasp of quantization tradeoffs
  • Experience with distributed inference across multi-GPU / multi-node environments
  • Fluency with GPU profiling tools (Nsight Systems, Nsight Compute, PyTorch Profiler)

Bonus points

  • C++ / CUDA proficiency
  • Diffusion / image-model inference optimization
  • Custom Triton kernel development
  • Contributions to open-source inference frameworks

Similar jobs