Member of Technical Staff (AI Inference Engineer)
Perplexity · Palo Alto, CA · 1 mo ago
On-siteEngineering$220k–$485k/yrFull-time
Who we're looking for
- Deep experience with GPU programming and performance work (CUDA, Triton, CUTLASS, or similar).
- Any other deep systems programming experience is a plus.
- You understand modern LLM architectures and are able to bring them up reliably in a production environment.
- You've built and operated production distributed systems under real load - ideally performance-critical ones.
- You can read a research paper on Monday, write a kernel on Wednesday, and debug a production incident on Friday.
- You own problems end-to-end.
- You are self-directed. You do well in fast-moving environments where the path forward isn't laid out for you.
- Good if you touched any of ML compilers and framework internals: PyTorch internals, torch.compile, custom operators.
- Distributed GPU communication: NCCL, NVLink, InfiniBand, RDMA libraries, model/tensor parallelism.
- Low-precision inference: INT8/FP8/FP4 quantization, mixed-precision serving.
- Profiling and debugging tools: Nsight Compute/Systems, CUDA-GDB, PTX/SASS analysis.
- Container orchestration: Kubernetes, GPU scheduling, autoscaling inference workloads.
Qualifications
- 3+ years of professional software engineering experience with meaningful work on ML inference or high-performance systems.
- Familiarity with at least one deep learning framework (PyTorch, JAX, TensorFlow).
- Understanding of GPU architectures (memory hierarchy, warp scheduling, tensor cores).
- Understanding of common LLM architectures and inference optimization techniques (e.g. quantization, speculative decoding, prefill-decode disaggregation).