Jobs · Engineering

Research And Development Engineer

Uber AI Solutions · United States · Yesterday
RemoteRemoteEngineeringPart-time

About the role

The Research Engineer (RE) Curator will design, implement, and review complex, multi-step ("mid-horizon") agentic tasks that simulate real-world challenges faced by Research Engineers. Each task will require 1-2 days of continuous effort to complete and will span multiple technical skills.

Responsibilities

  • Design and implement complex agentic tasks that simulate real-world challenges faced by Research Engineers.
  • Review and refine tasks to ensure they are challenging and effective.
  • Collaborate with researchers to provide feedback and iterate on task designs.
  • Work within a tight feedback loop to ensure tasks meet high standards.
  • Develop a new version of the RE-Bench evaluation benchmark by collecting high-quality, complex, "hard" tasks in a red-teaming setup.

Requirements

  • Must have an MSc or PhD in a STEM field, or equivalent practical experience in a research-heavy domain requiring data analysis and coding (e.g., computational sociology/humanities).
  • At least 1 year of experience in a research or research engineering role.
  • Basic proficiency in Python scripting.
  • Familiarity with Git and IDEs.
  • Fluent in English (written and verbal).
  • A perfectionist mindset, high attention to detail, creativity in task design, and ability to work independently and handle ambiguity.

Qualifications

  • Nice-to-have: Prior experience as a Research Scientist, Test Engineer, or Quality Reviewer.
  • Nice-to-have: Experience in Red Teaming for LLMs.

Skills

  • Python Scripting: Proficiency in writing and debugging Python code.
  • Development Infrastructure: Familiarity with version control systems (Git), Integrated Development Environments (IDEs), and basic software development workflows.
  • Agentic Coding: Conceptual understanding or experience using AI coding assistants (Gemini, Jetski) and understanding prompt engineering or agent workflows.
  • Software Quality: Clean code practices, readability, and basic debugging skills.
  • Machine Learning & Artificial Intelligence (Core): Foundational understanding of machine learning concepts, model training, and evaluation; familiarity with LLM capabilities, limitations, and evaluation techniques; basic understanding of reinforcement learning (RL); experience setting up, running, and analyzing ML experiments.
  • Data Science & Quantitative Analysis (Core): Heavy data analysis skills, including statistical correlation, data cleaning, and interpretation.
  • STEM Research & Experimental Methodology (Core): Strong background in experimental design, hypothesis testing, and rigorous evaluation; experience in STEM fields or Computational Humanities/Social Sciences requiring significant computational work.
  • Quality Assurance & Testing (Preferred / Plus): Experience designing test cases, quality review processes, and debugging complex systems.
  • AI Safety & Security (Preferred / Plus): Experience in identifying vulnerabilities, edge cases, or failure modes in LLMs.

Benefits

  • Tight feedback loop with researchers, including daily syncs.
  • Access to internal tools (Jetski, google3 codebase) and unreleased model checkpoints.

Pay

Compensation is commensurate with experience.

Schedule

35 hours per week.

Similar jobs