Jobs · Research

AI Vulnerability Expert - Fully Remote | Upto $90/hr

Mercor · United States · 5 days ago
RemoteRemoteResearch$60–$90/hrFull-time

Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors include Benchmark, General Catalyst, Peter Thiel, Adam D'Angelo, Larry Summers, and Jack Dorsey.

About the role

We are looking for an LLM Red Team Specialist to evaluate frontier AI models, identify vulnerabilities and failure modes, and help improve benchmark quality. This is a contract role with a commitment of 35 hours per week.

Responsibilities

  • Evaluate frontier AI models on coding, ML, and analysis tasks to identify vulnerabilities and failure modes.
  • Design complex tasks that challenge models and are fair for grading.
  • Document findings with clear evidence and reproducible steps.
  • Collaborate with task authors to close loopholes and improve grading.
  • Share insights with researchers to enhance benchmark quality.
  • Work independently and asynchronously to meet deadlines and improve AI model performance.

Requirements

  • MSc or PhD in a STEM field or equivalent experience.
  • 1+ years in research, research-engineering, security, or AI evaluation.
  • Experience identifying vulnerabilities in LLMs or ML systems.
  • Proficiency in Python and Git.
  • Familiarity with LLM capabilities and evaluation techniques.
  • Ability to work 35 hours per week.

Preferred Qualifications

  • Experience in AI training, model evaluation, or benchmark/task authoring.

Pay

$60–$90 per hour

Schedule

35 hours per week

For details about the interview process and platform information, visit https://talent.docs.mercor.com/welcome. For support, email support@mercor.com.

Similar jobs