INFERENCE OPTIMIZATION ENGINEER
Up Top · United States · 1 wk ago
RemoteRemoteEngineeringFull-time
About the role
This is a hybrid IC / leadership role reporting to the Head of Engineering. You'll own the technical strategy for inference performance at scale, then build and lead a dedicated Inference Optimization team. The mandate is simple to state and hard to execute: drive down latency, increase throughput, and cut cost-per-token for large-scale LLM inference across a modern GPU fleet.
What you'll do
- Own the end-to-end technical strategy for inference performance across the platform
- Recruit, build, and lead the Inference Optimization team
- Optimize GPU infrastructure across current and next-gen architectures (e.g. H200, B300)
- Improve latency, throughput, and cost-per-token for production LLM inference workloads
- Build reproducible benchmarking harnesses across inference engines such as vLLM and SGLang
- Tune inference routing and multivariate load-balancing algorithms
- Evaluate emerging optimization techniques — custom CUDA/Triton kernels, attention variants, quantization schemes, and compilation improvements
- Assess emerging inference hardware (FPGAs, ASICs, and custom silicon)
What we're looking for
- 8+ years in performance optimization or HPC, with deep GPU architecture and parallel-programming expertise
- 5+ years leading engineering teams
- Proficiency in Python, Rust, or Go
- Hands-on experience running production LLM inference engines at high volume
- Depth in modern inference optimization: continuous batching, PagedAttention / KV-cache management, speculative decoding, quantization, CUDA graphs, and torch.compile
- A strong grasp of quantization tradeoffs
- Experience with distributed inference across multi-GPU / multi-node environments
- Fluency with GPU profiling tools (Nsight Systems, Nsight Compute, PyTorch Profiler)
Bonus points
- C++ / CUDA proficiency
- Diffusion / image-model inference optimization
- Custom Triton kernel development
- Contributions to open-source inference frameworks