Research Scientist / Engineer – Reinforcement Learning Infrastructure
Luma · San Francisco Bay Area · 4 days ago
HybridOTHRFull-time
Responsibilities
- Design, build, and scale distributed RL post-training systems for large multimodal models — orchestrating trainer, rollout, environment, and reward workloads across thousands of GPUs
- Build and optimize high-throughput rollout generation, including efficient integration of inference engines (e.g. vLLM, SGLang) into the training loop, weight synchronization, and asynchronous / off-policy training schemes
- Design and implement RL environments for agentic and multi-step tasks — sandboxed code execution, tool use, computer use, and multimodal interaction — that are reproducible, hermetic, and scalable to millions of episodes
- Build reward infrastructure: verifiable / programmatic rewards, reward model serving, LLM-as-judge pipelines, and defenses against reward hacking
- Develop the evaluation, monitoring, and debugging tooling needed to keep large RL runs stable, diagnose convergence and throughput regressions, and understand model behavior mid-run
- Advance RL training efficiency and stability: sequence packing for long multi-turn trajectories, KV cache reuse across rollouts, curriculum and task sampling, and resource scheduling across heterogeneous training/inference workloads
Requirements
- Hands-on experience post-training LLMs with reinforcement learning (e.g. PPO / GRPO-family methods, RLHF, RLVR / RL from verifiable rewards) at meaningful scale
- Extensive experience with distributed PyTorch training and parallelization strategies (FSDP, Tensor / Pipeline / Expert Parallel) for foundation models
- Experience building RL environments, reward functions, verifiers, or evaluation harnesses for LLM agents — including sandboxed execution and multi-turn tool use
- Deep familiarity with RL post-training frameworks and their systems tradeoffs (e.g. veRL, OpenRLHF, TRL, Ray-based orchestration) and inference engines used for rollouts (vLLM, SGLang)
- Strong understanding of GPU clusters, networking, and communication libraries (NCCL, MPI), and how they behave under mixed training + inference workloads
- (Preferred) Experience running RL training across >100 GPUs, including asynchronous or disaggregated trainer/rollout architectures
- (Preferred) Experience with containerization and orchestration (Kubernetes, Ray) for large environment fleets and sandboxed workloads
Qualifications
- Research contributions in RL for LLMs — reasoning, agents, reward modeling, or long-horizon tasks — or open-source contributions to RL training frameworks