Member of Technical Staff - RL Training Framework
SpaceXAI · Palo Alto, CA · 1 wk ago
On-siteEngineering$180k–$440k/yrFull-time
About the role
The RL infrastructure team at SpaceXAI is seeking an engineer to assist in developing our Reinforcement Learning (RL) training framework.
Responsibilities
- Design and implement the systems supporting all RL workloads, ranging from small-scale experiments to full-scale production training runs.
- Profile, debug, and optimize the performance of end-to-end RL training processes.
- Enhance the scalability and observability of the RL stack.
BASIC QUALIFICATIONS
- Experience in building, debugging, and optimizing large-scale distributed systems.
- Comfortable working in unfamiliar areas and solving problems across different levels of the stack.
- Proficiency in Python, Jax, Rust, and/or C++.
PREFERRED SKILLS AND EXPERIENCE
- Experience with large-scale Language Model Training (LLM) infrastructure.
- Strong understanding of reinforcement learning techniques.
- Experience with RL numerics.
COMPENSATION AND BENEFITS
$180,000 - $440,000 USD
Base salary is part of our comprehensive total rewards package, which includes equity, medical, vision, dental coverage, a 401(k) retirement plan, disability insurance, life insurance, and various other benefits and perks. SpaceXAI is committed to an inclusive environment and offers a Recruitment Privacy Notice for details on data processing.