Staff Research Scientist / Reinforcement Learning
Oscar · San Jose, CA · 3 days ago
Research$250k–$350k/yrFull-time
About the role
A cutting-edge AI company is hiring a Staff AI Researcher to lead a small team while remaining hands-on across reinforcement learning and LLM post-training. This is a role where researchers take ownership over their work, collaborate closely with leadership, and help shape the technical direction of the organization.
Responsibilities
- Take ownership over their work and contribute individually to cutting-edge RL research and implementation
- Mentor junior researchers
- Build and improve RLHF pipelines
- Develop reward models
- Design post-training workflows using approaches like PPO, DPO, GRPO, KTO, and similar methods
- Work closely with evaluation frameworks to identify model weaknesses
- Build reinforcement learning environments
- Translate research into production-ready improvements
Requirements
- 5+ years of experience in AI Research, Applied Research, or Machine Learning
- PhD in Computer Science, Machine Learning, or a related technical field preferred
- 3+ years of experience fine-tuning LLMs using RLHF, PPO, DPO, GRPO, or similar post-training methods
- Experience building reinforcement learning environments, evaluation frameworks, and benchmark-driven systems
- Experience with modern post-training libraries such as TRL, OpenRLHF, veRL, SkyRL, or similar
- Strong publication record at ICML, NeurIPS, ICLR, ACL, or similar conferences preferred
- Experience mentoring researchers or leading technical initiatives while remaining hands-on
- Strong Python experience implementing machine learning research into production systems
Qualifications
- Hands-on experience fine-tuning frontier models
- Publishing RL research
- Leading technical initiatives in a collaborative research environment
Skills
- Experience with reinforcement learning and LLM post-training techniques
- Ability to build and improve RLHF pipelines, reward models, and post-training workflows
- Proficiency in modern post-training libraries
- Strong publication record in relevant fields
- Experience mentoring researchers or leading technical initiatives
- Strong Python skills for implementing machine learning research
Benefits
- Equity
- Paid Time Off
- Medical, dental, and vision coverage
- Opportunity to create lasting impact within a growing organization
- Opportunity to lead and grow a high-performing research team
- Hybrid working schedule