Runtime Engineer
MatX · Mountain View, CA · Yesterday
Information Technology$120k–$250k/yrFull-time
What You'll Do Here
- Build the host-side interface library — device memory management, DMA, streams and events, sync primitives — that every compiler-emitted program runs on top of
- Own and extend the executable format: the compiler→runtime contract, its versioning, the weight and quantization layouts that let compiler and runtime evolve independently
- Design the custom-kernel ABI — calling convention, sync semantics, lifecycle — and the host-side marshaling layer (DLPack, the buffer protocol, numpy) that gets Python tensors to the device
- Build Python bindings via PyO3, with a C-ABI shim as the alternative integration path for downstream consumers
- Build the LLM inference serving stack — paged KV cache, continuous batching, request scheduling, token streaming — and the cluster orchestration primitives underneath it
- Bring up interconnect topology from the host and own the failure-detection and clean-teardown path for stop-restructure-resume recovery across racks
- Design what the chip exposes to host-side profilers and debuggers — perf counters, traces, and the Python surfaces ML engineers actually use — and hit measurable performance targets on runtime overhead and serving throughput
Who You Are
- Strong experience in a systems programming language — Rust, C, C++, or Go — including memory management, allocator design, and FFI/ABI work
- Have built Python interop layers in production (PyO3, ctypes, pybind11, or equivalent C-ABI bridging)
- Have designed and maintained API or ABI contracts between teams — versioning, evolution, breaking-change discipline — not just consumed someone else's
- Hold hands-on with at least one accelerator programming model (CUDA, ROCm, oneAPI Level Zero, TPU, or comparable) — enough to reason about device memory, async execution, and kernel launch
- ML-systems literate — comfortable with the training and inference loop, what collectives do, what a tensor layout is. Research depth not required.
Bonus Points
- If You Have LLM inference internals — vLLM, TensorRT-LLM, or SGLang (paged attention, scheduler design)
- Rust at depth, including proc macros, unsafe with soundness reasoning, and complex lifetime/trait work
- Custom allocator design (slab, paged, arena) or other low-level memory work
- ML framework integration experience (PyTorch custom backends, JAX/XLA, ONNX runtime)
- Profiler or tracing infrastructure work (perfetto, Nsight, or a custom stack)
- Driver-adjacent or kernel-bypass work, or prior new-silicon bring-up
Compensation
The US base salary for this full-time position is determined based on a variety of factors including role, experience, location, job related skills, and relevant education and training.
What We Offer
- A Stake in our success
- A flexible cash equity compensation mix that fits your needs
- Health & Wellness
- Time To Recharge
- Support to Parents
- Learning & Development
- Financial Wellbeing
- Remote Perks
MatX E[x]tras
$50 per month to use on the perks you care about most
As an Equal Opportunity Employer
We do not discriminate on the basis of race, color, religion, creed, national origin or ancestry, sex, gender, gender identity, gender expression, sexual orientation, age, physical or mental disability, medical condition, marital/domestic partner status, military and veteran status, genetic information or any other legally recognized protected basis under federal, state or local laws, regulations or ordinances.