Jobs · Engineering · California

Sr. Inference Optimization Engineer (local / edge runtime)

Intel · Santa Clara, CA · 1 mo ago
HybridEngineering$195k–$361k/yrFull-time

About the role

Make models fast on the hardware people actually own. Optimize inference engines (llama.cpp, vLLM) for constrained local and edge environments — GPU/iGPUs, Vulkan backends — not datacenter H100 environment, mostly PC/edge.

Responsibilities

  • Profile and optimize local inference (llama.cpp-vulkan and vLLM) for latency, throughput, and memory on edge hardware
  • Tune KV cache, continuous batching, and scheduling for interactive agent workloads
  • Drive quantization strategy (GGUF / AWQ / GPTQ) and validate quality impact with the Post-Training team
  • Cut CPU overhead and improve engine startup, model load, and lifecycle (start / stop / health)
  • Benchmark across hardware tiers and publish honest performance comparisons
  • Upstream fixes and patches to open-source engines where it helps us

Qualifications

  • BS/MS in CS, EE, Math or related STEM field
  • 8+ years software development background
  • Strong in C++ and/or Python; comfortable reading systems-level code
  • Experience with LLM inference. (attention, KV cache, decoding)
  • Experience profiling and optimizing real performance problems (CPU or GPU)
  • Linux, build systems, and low-level debugging expertise

Requirements

  • Hands-on with llama.cpp, vLLM, ggml, or similar engines
  • Experience with GPU / accelerator programming (Vulkan, CUDA, SYCL, Metal)
  • Familiarity with quantization formats and their quality trade-offs
  • Open-source contributions to inference engines

Benefits

Intel offers a total compensation package that ranks among the best in the industry. It consists of competitive pay, stock bonuses, and benefit programs which include health, retirement, and vacation.

Pay

$195,200.00 - 361,200.00 USD

Schedule

Shift 1 (United States of America)

Similar jobs