Jobs · Research · California

Staff Research Scientist / Reinforcement Learning

Oscar · San Jose, CA · 3 days ago
Research$250k–$350k/yrFull-time

About the role

A cutting-edge AI company is hiring a Staff AI Researcher to lead a small team while remaining hands-on across reinforcement learning and LLM post-training. This is a role where researchers take ownership over their work, collaborate closely with leadership, and help shape the technical direction of the organization.

Responsibilities

  • Take ownership over their work and contribute individually to cutting-edge RL research and implementation
  • Mentor junior researchers
  • Build and improve RLHF pipelines
  • Develop reward models
  • Design post-training workflows using approaches like PPO, DPO, GRPO, KTO, and similar methods
  • Work closely with evaluation frameworks to identify model weaknesses
  • Build reinforcement learning environments
  • Translate research into production-ready improvements

Requirements

  • 5+ years of experience in AI Research, Applied Research, or Machine Learning
  • PhD in Computer Science, Machine Learning, or a related technical field preferred
  • 3+ years of experience fine-tuning LLMs using RLHF, PPO, DPO, GRPO, or similar post-training methods
  • Experience building reinforcement learning environments, evaluation frameworks, and benchmark-driven systems
  • Experience with modern post-training libraries such as TRL, OpenRLHF, veRL, SkyRL, or similar
  • Strong publication record at ICML, NeurIPS, ICLR, ACL, or similar conferences preferred
  • Experience mentoring researchers or leading technical initiatives while remaining hands-on
  • Strong Python experience implementing machine learning research into production systems

Qualifications

  • Hands-on experience fine-tuning frontier models
  • Publishing RL research
  • Leading technical initiatives in a collaborative research environment

Skills

  • Experience with reinforcement learning and LLM post-training techniques
  • Ability to build and improve RLHF pipelines, reward models, and post-training workflows
  • Proficiency in modern post-training libraries
  • Strong publication record in relevant fields
  • Experience mentoring researchers or leading technical initiatives
  • Strong Python skills for implementing machine learning research

Benefits

  • Equity
  • Paid Time Off
  • Medical, dental, and vision coverage
  • Opportunity to create lasting impact within a growing organization
  • Opportunity to lead and grow a high-performing research team
  • Hybrid working schedule

Similar jobs