Jobs · Engineering · California

Principal LLM Inference Engineer

d-Matrix · Santa Clara, CA · 1 mo ago
HybridEngineeringFull-time

Responsibilities

  • Identify and prototype emerging LLM inference use cases suited to heterogeneous hardware deployments.
  • Build compelling proof-of-concept systems that demonstrate D-Matrix capabilities to customers, partners, and internal stakeholders.
  • Develop and tune custom kernels and operator-level optimizations to maximize throughput and minimize latency.
  • Drive quantization, sparsity, and batching strategies tailored to D-Matrix computational model.
  • Build and maintain inference runtimes, serving frameworks, and evaluation tooling.
  • Contribute to distributed inference systems: tensor/pipeline parallelism, disaggregated prefill/decode, KV-cache management.
  • Work closely with hardware architects to provide firmware and compiler teams with actionable inference workload insights.
  • Partner with product and business development to translate POCs into customer-facing demonstrations.
  • Contribute to technical publications, whitepapers, and open-source projects that advance D-Matrix visibility.

Requirements

  • Bachelor’s degree in Computer Science, Electrical Engineering, or a related field, and 10+ years of relevant engineering experience; or equivalent demonstrated experience.
  • Master’s or PhD in Computer Science, Electrical Engineering, or a related field preferred, with 6+ years of relevant industry experience.
  • Strong proficiency in Python and C/C++.
  • Hands-on experience optimizing LLM inference — attention kernels, KV cache, batching strategies, quantization (INT8/FP8/INT4).
  • Experience with at least one major inference framework (vLLM, SGLang, TensorRT-LLM, ONNX Runtime, or similar) at a contributor level.
  • Familiarity with GPU kernel programming (CUDA/Triton) and performance profiling tools.

Preferred Qualifications

  • Experience with heterogeneous compute deployments — scheduling inference workloads across dissimilar hardware (accelerators, CPUs, GPUs).
  • Familiarity with custom silicon or ASIC-based inference (beyond GPU-only environments).
  • Experience with distributed inference: tensor parallelism, pipeline parallelism, disaggregated serving.
  • Contributions to open-source inference or ML systems projects.
  • Experience with production inference serving at scale (latency SLOs, continuous batching, multi-model serving).
  • Familiarity with speculative decoding, mixture-of-experts routing, or long-context serving techniques.
  • Working familiarity with the material in the JAX Scaling Book or equivalent systems-level understanding of modern LLM training and inference.

Benefits

Competitive compensation, equity, and benefits in Santa Clara, CA.

Pay

TBD

Schedule

TBD

Similar jobs

Inference Engineer

CartesiaSan Francisco, CA· 3 wk ago
Engineeringapply on jobs.ashbyhq.com