AI Research Scientist - Infrastructure Engineer, Reinforcement Learning
AMD · Santa Clara, CA · 2 wk ago
HybridEngineeringFull-time
About the Role
We are hiring an AI Research Scientist - Infrastructure Engineer, Reinforcement Learning, to own reinforcement learning infrastructure at scale—including distributed policy and value training, rollout generation, logging, checkpointing, and researcher-facing APIs across large GPU fleets. You make RL scientists productive by improving throughput, fault tolerance, reproducibility, and observability—turning fragile notebooks into reliable systems that RSI, generalizing-HW, and RL research programs depend on.
Responsibilities
- Design and implement distributed RL training stacks (data parallel, pipeline parallel, or hybrid) integrated with AMD’s schedulers and storage
- Build high-throughput rollout workers, trajectory stores, and reward computation pipelines with versioning and audit trails
- Instrument jobs for debugging (NaNs, stragglers, OOMs), implement autoscaling and preemption-safe checkpointing
- Collaborate with research scientists on experiment templates, hyperparameter sweeps, and safe promotion paths from research to wider team use
- Drive reliability: on-call rotations, runbooks, and postmortems for infra incidents affecting RL training
Qualifications
- Strong systems track record in machine learning (ML) platforms with deep systems expertise and demonstrated technical impact
- Deep experience with PyTorch (or JAX), NCCL/MPI-style distributed training, and GPU cluster orchestration
- Prior ownership of RL training infra, LLM post-training pipelines, or large-scale experiment management
- Proficiency in C++/Python performance tuning, I/O optimization, and containerized workloads
Academic Credentials
Bachelor’s degree required; Master’s or PhD preferred in Computer Science for research-heavy collaboration depth.
Benefits
AMD offers a comprehensive benefits package. For more details, see AMD benefits at a glance.