Jobs · Information Technology · California

Helix AI Engineer, Training Performance

Jobera.com · San Jose, CA · 4 wk ago
Information Technology$200k–$400k/yrFull-time

About the role

Figure is an AI robotics company developing autonomous general-purpose humanoid robots with human-level intelligence. Our goal is to deploy humanoid robots at a global scale, performing a variety of tasks in both home and commercial markets. As part of our Helix team, you will focus on improving distributed training frameworks for large-scale model training, optimizing GPU kernels, evaluating accelerator types, and co-designing models to maximize hardware utilization.

Responsibilities

  • Optimize training performance for 100B+ parameter models across 100k+ GPUs.
  • Collaborate on accelerator choice, cluster topology, scheduling, and hardware procurement to inform future scaling.
  • Write and optimize custom kernels (Triton/CUDA).
  • Build tooling and dashboards for continuous performance monitoring, regression detection, and root-cause analysis.
  • Optimize data loading and preprocessing pipelines to prevent I/O bottlenecks.
  • Improve checkpointing, fault tolerance, and elastic restart for rapid recovery from node failures.
  • Partner with researchers to co-design performant model architectures and training recipes (e.g., activation checkpointing, mixed precision, sequence packing).
  • Extend and contribute to kernel compilers (e.g., Triton, Gluon) to improve iteration speed and enable targeting of custom/non-NVIDIA accelerators.
  • Build agentic systems to automatically generate, benchmark, and iterate on custom kernels.
  • Evaluate emerging accelerator architectures (AMD, TPU, SRAM-based ASICs, etc.) and lead proof-of-concept ports/benchmarks.
  • Explore model/data parallelisms (FSDP, context parallel, expert parallel, etc.) to determine optimal configurations per model size.

Requirements

  • Bachelor's or Master's degree in Computer Science, Computer/Electrical Engineering, or a related field.
  • 3+ years in AI performance engineering, with experience leading large-scale performance improvement projects.
  • Deep understanding of GPU architecture and performance characteristics (memory bandwidth, compute-bound vs. memory-bound ops, occupancy).
  • Proficiency with profiling tools (Nsight Systems/Compute, PyTorch Profiler, HTA, or similar) and translating traces into optimizations.
  • Solid grasp of collective communication (NCCL) and modern networking (RDMA, NVLink, InfiniBand/RoCE, topology-aware placement).
  • Strong Python and CUDA/C++ skills; comfortable modifying framework internals.
  • Experience debugging performance regressions and instability at scale (stragglers, hangs, OOMs, numerical divergence).
  • Experience defining and reasoning about hardware-efficiency metrics (MFU/HFU) to drive optimization priorities.
  • Experience with heterogeneous or multi-datacenter training setups and cross-cluster orchestration.
  • Contributions to open-source ML systems projects (PyTorch, Megatron-LM, vLLM, DeepSpeed, JAX, etc.).
  • Exposure to non-NVIDIA accelerators (AMD GPUs, TPU/Trainium/Inferentia, or custom silicon) and heterogeneous fleet management.

Pay

The US base salary range for this full-time position is $200,000 - $400,000 annually. The final pay offered may vary based on job-related knowledge, skills, and experience. Additional compensation components may be included in the total package.

Similar jobs

Helix AI Engineer, iOS

FigureSan Jose, CA· 1 mo ago
Engineering$200k–$400k/yrapply on job-boards.greenhouse.io