Research Engineer - AI-Optimized Inference
Touring Capital · San Francisco Bay Area · Yesterday
Engineering$185k–$455k/yrFull-time
About the role
The role involves optimizing inference stacks to achieve at least 50% faster throughput compared to existing libraries, measured on real-world workloads. The focus is on rewriting kernels and their surrounding layers to improve performance on AI accelerators.
Responsibilities
- Own the optimization loop that refines inference kernels until they meet the 50% performance target.
- Generate and evaluate candidate kernel rewrites on real inference workloads, keeping only those that improve performance.
- Ensure evaluations are fair and not easily gameable, focusing on throughput improvements rather than misleading benchmarks.
- Work on kernels including matmul, attention, normalization, and collective operations, ensuring they align with the accelerator's peak performance.
- Develop and maintain the evaluation harness that drives the optimization loop, ensuring that each improvement is validated against a reference implementation.
- Collaborate with the Infinity Kernel Registry to reuse successful optimizations across different models and chips.
- Ensure all optimized kernels are correct and produce accurate results, not just faster versions of incorrect outputs.
Requirements
- Experience with performance engineering on real inference or accelerator code.
- Familiarity with search- or evolution-based optimization techniques, such as AlphaEvolve.
- A solid understanding of inference internals, including kernels, batching, KV cache, scheduling, and throughput bottlenecks.
- Strong measurement discipline that avoids being swayed by benchmarks designed to inflate performance numbers.
- Proficiency in Python and a systems language, with hands-on experience in CUDA, ROCm/HIP, Triton, or similar environments.
- Optional skills include written high-performance attention or matmul kernels, contributions to inference serving stacks, and experience with code evolution tools.