Jobs · Research

AI Safety Expert - Red Team

Mercor · United States · 1 wk ago
RemoteRemoteResearch$20–$22/hrPart-time

Role Responsibilities

  • Red team conversational AI models and agents to identify jailbreaks, prompt injections, misuse cases, and bias exploitation.
  • Generate high-quality human data by annotating failures, classifying vulnerabilities, and flagging systemic risks.
  • Apply structure by following taxonomies, benchmarks, and playbooks to maintain consistent testing.
  • Document reproducibly to produce reports, datasets, and attack cases that customers can act on.
  • Work independently and asynchronously to meet deadlines while improving AI model performance.

Qualifications

  • Must-Have:
    • Fluent Language Skills: Required: English & Urdu. Native fluency in English and Urdu is required.
    • Prior red teaming experience in AI adversarial work, cybersecurity, or socio-technical probing.
    • Strong communication skills to explain risks clearly to technical and non-technical stakeholders.
    • Ability to thrive on moving across projects and customers.
  • Preferred Experience:
    • In Adversarial ML: jailbreak datasets, prompt injection, RLHF/DPO attacks, model extraction.
    • In Cybersecurity: penetration testing, exploit development, reverse engineering.
    • In socio-technical risk: harassment/disinfo probing, abuse analysis, conversational AI testing.
    • In Creative probing skills: psychology, acting, writing for unconventional adversarial thinking.

Similar jobs