Machine Learning Research Scientist, Evaluations
Scale AI · Seattle, WA · Today
OTHR$181k–$226k/yrFull-time
About the role
Scale works with the industry's leading AI labs to provide high quality data and accelerate progress in GenAI research. We are looking for Research Scientists and Research Engineers with expertise in LLM post-training (SFT, RLHF, reward modeling) and evaluation.
Responsibilities
- Analyze model behavior to identify, characterize, and diagnose failure modes in frontier LLMs and Agents.
- You’ll identify everything from capability gaps and reasoning errors to robustness and alignment issues, all focusing on RCA.
- Design and build benchmarks and evaluation methods that measure LLM capabilities in both text and multimodal modalities.
- Apply post-training expertise (SFT, RLHF, reward modeling) to connect observed failures to the data and training interventions that address them.
- Develop rigorous evaluations and diagnostic methods that reveal where frontier models fail and why.
- Collaborate with researchers and engineers to define best practices in evaluation-driven AI development.
- Partner with top foundation model labs to translate failure analysis into technical and strategic input on the next generation of generative AI models.
Requirements
- Ph.D. or Master's degree in Computer Science, Machine Learning, AI, or a related field.
- Deep understanding of deep learning, reinforcement learning, and large-scale model fine-tuning.
- Experience with post-training techniques such as RLHF, preference modeling, or instruction tuning, and with LLM evaluation or benchmark development.
- Excellent written and verbal communication skills.
- Published research in areas of machine learning at major conferences (NeurIPS, ICML, ICLR, ACL, EMNLP, CVPR, etc.) and/or journals.
- Previous experience in a customer facing role.
Qualifications
- Strong analytical and problem-solving skills.
- Ability to work independently and as part of a team.
- Experience with natural language processing and multimodal analysis.
- Knowledge of ethical considerations in AI development.
Skills
- Expertise in LLM post-training techniques (SFT, RLHF, reward modeling).
- Experience with benchmark development and evaluation methods.
- Strong communication and collaboration skills.
- Ability to publish research in top-tier AI conferences.
Benefits
- Comprehensive health, dental and vision coverage.
- Roth IRA matching.
- Learning and development stipend.
- Generous PTO.
Pay
- The base salary range for this full-time position in the locations of San Francisco, New York, Seattle is: $180,600—$225,750 USD.
Schedule
- This role is full-time.
Location
This role is based in the San Francisco, New York, or Seattle area.