Helix AI Engineer, Training Performance
Jobera.com · San Jose, CA · 4 wk ago
Information Technology$200k–$400k/yrFull-time
About the role
Figure is an AI robotics company developing autonomous general-purpose humanoid robots with human-level intelligence. Our goal is to deploy humanoid robots at a global scale, performing a variety of tasks in both home and commercial markets. As part of our Helix team, you will focus on improving distributed training frameworks for large-scale model training, optimizing GPU kernels, evaluating accelerator types, and co-designing models to maximize hardware utilization.
Responsibilities
- Optimize training performance for 100B+ parameter models across 100k+ GPUs.
- Collaborate on accelerator choice, cluster topology, scheduling, and hardware procurement to inform future scaling.
- Write and optimize custom kernels (Triton/CUDA).
- Build tooling and dashboards for continuous performance monitoring, regression detection, and root-cause analysis.
- Optimize data loading and preprocessing pipelines to prevent I/O bottlenecks.
- Improve checkpointing, fault tolerance, and elastic restart for rapid recovery from node failures.
- Partner with researchers to co-design performant model architectures and training recipes (e.g., activation checkpointing, mixed precision, sequence packing).
- Extend and contribute to kernel compilers (e.g., Triton, Gluon) to improve iteration speed and enable targeting of custom/non-NVIDIA accelerators.
- Build agentic systems to automatically generate, benchmark, and iterate on custom kernels.
- Evaluate emerging accelerator architectures (AMD, TPU, SRAM-based ASICs, etc.) and lead proof-of-concept ports/benchmarks.
- Explore model/data parallelisms (FSDP, context parallel, expert parallel, etc.) to determine optimal configurations per model size.
Requirements
- Bachelor's or Master's degree in Computer Science, Computer/Electrical Engineering, or a related field.
- 3+ years in AI performance engineering, with experience leading large-scale performance improvement projects.
- Deep understanding of GPU architecture and performance characteristics (memory bandwidth, compute-bound vs. memory-bound ops, occupancy).
- Proficiency with profiling tools (Nsight Systems/Compute, PyTorch Profiler, HTA, or similar) and translating traces into optimizations.
- Solid grasp of collective communication (NCCL) and modern networking (RDMA, NVLink, InfiniBand/RoCE, topology-aware placement).
- Strong Python and CUDA/C++ skills; comfortable modifying framework internals.
- Experience debugging performance regressions and instability at scale (stragglers, hangs, OOMs, numerical divergence).
- Experience defining and reasoning about hardware-efficiency metrics (MFU/HFU) to drive optimization priorities.
- Experience with heterogeneous or multi-datacenter training setups and cross-cluster orchestration.
- Contributions to open-source ML systems projects (PyTorch, Megatron-LM, vLLM, DeepSpeed, JAX, etc.).
- Exposure to non-NVIDIA accelerators (AMD GPUs, TPU/Trainium/Inferentia, or custom silicon) and heterogeneous fleet management.
Pay
The US base salary range for this full-time position is $200,000 - $400,000 annually. The final pay offered may vary based on job-related knowledge, skills, and experience. Additional compensation components may be included in the total package.