AI Content Red Team Analyst - Trust and Safety
About the role
The Trust & Safety (T&S) GenAI & Emerging Product team empowers the development of GenAI models and applications by building a world-class safety, testing, and risk management system to ensure responsible launches. The AI Content Red Team, part of the T&S GenAI and Emerging Products pillar, conducts unstructured adversarial testing of TikTok's generative AI products and models to uncover emerging risks alongside structured evaluations. This team combines attacker-minded testing, risk discovery, and operational feedback loops to inform product decisions, policy development, mitigations, and evaluation strategy. The role involves probing models and product experiences across modalities, use cases, and abuse patterns to identify failure modes and stress-test safeguards.
Responsibilities
- Conduct structured adversarial testing on AI models, features, and policies to identify vulnerabilities and emerging risks.
- Explore product behavior across contexts and user journeys to identify model failure modes not captured in standard evaluations.
- Investigate jailbreaks, evasions, prompt-based attacks, and other adversarial techniques relevant to content safety.
- Document findings clearly, including risk descriptions, reproduction steps, severity assessments, and mitigation recommendations.
- Partner with cross-functional stakeholders (policy, product, business teams) to ensure mitigation validation and root cause closure.
- Support development of testing playbooks, taxonomies, and internal knowledge bases.
- Stay updated on emerging adversarial trends (e.g., deepfakes, multimodal manipulation, coordinated abuse) and shifts in the external risk landscape.
Requirements
This role may expose you to harmful content related to bullying, hate speech, child safety, depictions of harm to self/others, and harm to animals. TikTok provides comprehensive wellbeing programs to support employees in this psychologically demanding work.
Qualifications
Minimum Qualifications:
- Minimum 3 years of experience in Trust & Safety, cybersecurity, risk/adversarial testing, or related fields.
- Experience with prompt testing, jailbreak analysis, LLM evaluation, or adversarial QA.
- Familiarity with AI safety risks (jailbreaks, hallucinations, bias, misuse patterns).
- Strong interest in GenAI safety and how AI systems can be compromised under adversarial conditions.
- Demonstrated ability to independently investigate ambiguous problems, identify non-obvious failure modes, and produce clear, evidence-based conclusions.
- Ability to manage multiple priorities and collaborate effectively with cross-functional teams.
Preferred Qualifications:
- Experience working with agentic AI tools to scale impact, including building/operating AI tools for process efficiency.
About TikTok
TikTok is the leading destination for short-form mobile video, with a mission to inspire creativity and bring joy. Headquartered in Los Angeles and Singapore, TikTok has global offices in cities including New York, London, Dublin, Paris, Berlin, Dubai, Jakarta, Seoul, and Tokyo.
Why Join Us
TikTok fosters an inclusive environment where diverse voices are celebrated. We embrace curiosity, humility, and resilience, working collaboratively to innovate and create value for our communities. Every challenge is an opportunity to learn and grow together.
Pay
The base salary range for this position is $121,600 – $272,000 annually, depending on qualifications, skills, experience, and location. Additional compensation may include discretionary bonuses, incentives, and restricted stock units. Benefits include medical, dental, and vision insurance; 401(k) with company match; paid parental leave; short-term and long-term disability coverage; life insurance; wellbeing benefits; 10 paid holidays; 10 paid sick days; and 17 days of paid personal time (prorated upon hire).