Sr. Inference Optimization Engineer (local / edge runtime)
Intel · Santa Clara, CA · 1 mo ago
HybridEngineering$195k–$361k/yrFull-time
About the role
Make models fast on the hardware people actually own. Optimize inference engines (llama.cpp, vLLM) for constrained local and edge environments — GPU/iGPUs, Vulkan backends — not datacenter H100 environment, mostly PC/edge.
Responsibilities
- Profile and optimize local inference (llama.cpp-vulkan and vLLM) for latency, throughput, and memory on edge hardware
- Tune KV cache, continuous batching, and scheduling for interactive agent workloads
- Drive quantization strategy (GGUF / AWQ / GPTQ) and validate quality impact with the Post-Training team
- Cut CPU overhead and improve engine startup, model load, and lifecycle (start / stop / health)
- Benchmark across hardware tiers and publish honest performance comparisons
- Upstream fixes and patches to open-source engines where it helps us
Qualifications
- BS/MS in CS, EE, Math or related STEM field
- 8+ years software development background
- Strong in C++ and/or Python; comfortable reading systems-level code
- Experience with LLM inference. (attention, KV cache, decoding)
- Experience profiling and optimizing real performance problems (CPU or GPU)
- Linux, build systems, and low-level debugging expertise
Requirements
- Hands-on with llama.cpp, vLLM, ggml, or similar engines
- Experience with GPU / accelerator programming (Vulkan, CUDA, SYCL, Metal)
- Familiarity with quantization formats and their quality trade-offs
- Open-source contributions to inference engines
Benefits
Intel offers a total compensation package that ranks among the best in the industry. It consists of competitive pay, stock bonuses, and benefit programs which include health, retirement, and vacation.
Pay
$195,200.00 - 361,200.00 USD
Schedule
Shift 1 (United States of America)