Jobs · Oregon

Sr. Inference Optimization Engineer (local / edge runtime)

Intel · Hillsboro, OR · 1 mo ago
Hybrid$195k–$361k/yrFull-time

About the role

Make models fast on the hardware people actually own. You optimize inference engines (llama.cpp, vLLM) for constrained local and edge environments — GPU/iGPUs, Vulkan backends — not datacenter H100 environment, mostly PC/edge. KV cache, batching, quantization, scheduling, and CPU-overhead reduction are your daily tools. This is the rare skill that makes a hybrid, low-cost agent product viable.

At Intel, our journey is to transform AI into something safer, more trustworthy, and respectful of human privacy by design. We believe transformative AI should have a positive impact on people—powerful in capability, yet honest about its limits and protective of the data and resources it touches. To get there, we build agentic AI that combines the best of local and cloud intelligence — private, affordable, and sustainable by design. Small, efficient models run directly on the user's machine (AI PC, edge, on-prem, and beyond), keeping data private and token costs low, while powerful cloud models handle the hardest work: planning, reasoning, and complex problem-solving.

Responsibilities

  • Profile and optimize local inference (llama.cpp-vulkan and vLLM) for latency, throughput, and memory on edge hardware
  • Tune KV cache, continuous batching, and scheduling for interactive agent workloads
  • Drive quantization strategy (GGUF / AWQ / GPTQ) and validate quality impact with the Post-Training team
  • Cut CPU overhead and improve engine startup, model load, and lifecycle (start / stop / health)
  • Benchmark across hardware tiers and publish honest performance comparisons
  • Upstream fixes and patches to open-source engines where it helps us

What you’ll learn / grow into

  • The internals of modern inference engines and where the milliseconds actually go
  • Hardware-aware optimization across iGPU / CPU paths (Vulkan, SYCL, oneAPI, CUDA where relevant)
  • The quality-vs-speed-vs-memory trade space for small models
  • Interest in local / edge AI and squeezing hardware

Requirements

Minimum qualifications:

  • BS/MS in CS, EE, Math or related STEM field
  • 8+ years software development background
  • Strong in C++ and/or Python; comfortable reading systems-level code
  • Experience with LLM inference (attention, KV cache, decoding)
  • Experience profiling and optimizing real performance problems (CPU or GPU) and can prove the speedup
  • Linux, build systems, and low-level debugging expertise

Preferred qualifications:

  • Hands-on with llama.cpp, vLLM, ggml, or similar engines
  • Experience with GPU / accelerator programming (Vulkan, CUDA, SYCL, Metal) or SIMD / CPU kernels
  • Familiarity with quantization formats and their quality trade-offs
  • Open-source contributions to inference engines

Requirements listed would be obtained through a combination of industry relevant job experience, internship experiences and/or schoolwork/classes/research.

Benefits

We offer a total compensation package that ranks among the best in the industry. It consists of competitive pay, stock bonuses, and benefit programs which include health, retirement, and vacation. Find out more about the benefits of working at Intel.

Pay

Annual Salary Range for jobs which could be performed in the US: $195,200.00 - 361,200.00 USD. The range displayed reflects the minimum and maximum target compensation for the position across all US locations. Within the range, individual pay is determined by work location and additional factors, including job-related skills, experience, and relevant education or training.

Schedule

This role will be eligible for our hybrid work model which allows employees to split their time between working on-site at their assigned Intel site and off-site.

Similar jobs