Research And Development Engineer
Uber AI Solutions · United States · Yesterday
RemoteRemoteEngineeringPart-time
About the role
The Research Engineer (RE) Curator will design, implement, and review complex, multi-step ("mid-horizon") agentic tasks that simulate real-world challenges faced by Research Engineers. Each task will require 1-2 days of continuous effort to complete and will span multiple technical skills.
Responsibilities
- Design and implement complex agentic tasks that simulate real-world challenges faced by Research Engineers.
- Review and refine tasks to ensure they are challenging and effective.
- Collaborate with researchers to provide feedback and iterate on task designs.
- Work within a tight feedback loop to ensure tasks meet high standards.
- Develop a new version of the RE-Bench evaluation benchmark by collecting high-quality, complex, "hard" tasks in a red-teaming setup.
Requirements
- Must have an MSc or PhD in a STEM field, or equivalent practical experience in a research-heavy domain requiring data analysis and coding (e.g., computational sociology/humanities).
- At least 1 year of experience in a research or research engineering role.
- Basic proficiency in Python scripting.
- Familiarity with Git and IDEs.
- Fluent in English (written and verbal).
- A perfectionist mindset, high attention to detail, creativity in task design, and ability to work independently and handle ambiguity.
Qualifications
- Nice-to-have: Prior experience as a Research Scientist, Test Engineer, or Quality Reviewer.
- Nice-to-have: Experience in Red Teaming for LLMs.
Skills
- Python Scripting: Proficiency in writing and debugging Python code.
- Development Infrastructure: Familiarity with version control systems (Git), Integrated Development Environments (IDEs), and basic software development workflows.
- Agentic Coding: Conceptual understanding or experience using AI coding assistants (Gemini, Jetski) and understanding prompt engineering or agent workflows.
- Software Quality: Clean code practices, readability, and basic debugging skills.
- Machine Learning & Artificial Intelligence (Core): Foundational understanding of machine learning concepts, model training, and evaluation; familiarity with LLM capabilities, limitations, and evaluation techniques; basic understanding of reinforcement learning (RL); experience setting up, running, and analyzing ML experiments.
- Data Science & Quantitative Analysis (Core): Heavy data analysis skills, including statistical correlation, data cleaning, and interpretation.
- STEM Research & Experimental Methodology (Core): Strong background in experimental design, hypothesis testing, and rigorous evaluation; experience in STEM fields or Computational Humanities/Social Sciences requiring significant computational work.
- Quality Assurance & Testing (Preferred / Plus): Experience designing test cases, quality review processes, and debugging complex systems.
- AI Safety & Security (Preferred / Plus): Experience in identifying vulnerabilities, edge cases, or failure modes in LLMs.
Benefits
- Tight feedback loop with researchers, including daily syncs.
- Access to internal tools (Jetski, google3 codebase) and unreleased model checkpoints.
Pay
Compensation is commensurate with experience.
Schedule
35 hours per week.