AI Safety Expert - Red Team
Mercor · United States · 1 wk ago
RemoteRemoteResearch$20–$22/hrPart-time
Role Responsibilities
- Red team conversational AI models and agents to identify jailbreaks, prompt injections, misuse cases, and bias exploitation.
- Generate high-quality human data by annotating failures, classifying vulnerabilities, and flagging systemic risks.
- Apply structure by following taxonomies, benchmarks, and playbooks to maintain consistent testing.
- Document reproducibly to produce reports, datasets, and attack cases that customers can act on.
- Work independently and asynchronously to meet deadlines while improving AI model performance.
Qualifications
- Must-Have:
- Fluent Language Skills: Required: English & Urdu. Native fluency in English and Urdu is required.
- Prior red teaming experience in AI adversarial work, cybersecurity, or socio-technical probing.
- Strong communication skills to explain risks clearly to technical and non-technical stakeholders.
- Ability to thrive on moving across projects and customers.
- Preferred Experience:
- In Adversarial ML: jailbreak datasets, prompt injection, RLHF/DPO attacks, model extraction.
- In Cybersecurity: penetration testing, exploit development, reverse engineering.
- In socio-technical risk: harassment/disinfo probing, abuse analysis, conversational AI testing.
- In Creative probing skills: psychology, acting, writing for unconventional adversarial thinking.