Jobs · Management · California

AI Content Red Team Analyst - Trust and Safety

TikTok · San Jose, CA · 1 wk ago
On-siteManagement$109k–$288k/yrFull-time

About the role

The Trust & Safety (T&S) GenAI & Emerging Product team empowers the development of GenAI models and applications by building a world-class safety, testing, and risk management system to ensure responsible launches. The AI Content Red Team, part of this pillar, conducts unstructured adversarial testing of Bytedance's generative AI products and models to uncover emerging risks alongside structured evaluations. This team combines attacker-minded testing, risk discovery, and operational feedback to inform product decisions, policy development, mitigations, and evaluation strategy. Work involves probing models and product experiences across modalities, use cases, and abuse patterns to identify failure modes and stress-test safeguards.

Responsibilities

  • Conduct structured adversarial testing on AI models, features, and policies to identify vulnerabilities and emerging risks.
  • Explore product behavior across contexts and user journeys to identify model failure modes not captured in standard evaluations.
  • Investigate jailbreaks, evasions, prompt-based attacks, and other adversarial techniques relevant to content safety.
  • Document findings clearly, including risk descriptions, reproduction steps, severity assessments, and mitigation recommendations.
  • Partner with cross-functional stakeholders (policy, product, business teams) to ensure mitigation validation and root cause closure.
  • Support development of testing playbooks, taxonomies, and internal knowledge bases.
  • Stay updated on emerging adversarial trends (e.g., deepfakes, multimodal manipulation, coordinated abuse) and shifts in the external risk landscape.

Requirements

  • Minimum 3 years of experience in Trust & Safety, cybersecurity, risk/adversarial testing, or related fields.
  • Experience with prompt testing, jailbreak analysis, LLM evaluation, or adversarial QA.
  • Familiarity with AI safety risks (jailbreaks, hallucinations, bias, misuse patterns).
  • Strong interest in GenAI safety and understanding of how AI systems can be compromised under adversarial conditions.
  • Demonstrated ability to independently investigate ambiguous problems, identify non-obvious failure modes, and produce clear, evidence-based conclusions.
  • Ability to manage multiple priorities and collaborate effectively with cross-functional teams.

Preferred Qualifications

  • Experience working with agentic AI tools to scale impact, including building/operating AI tools to improve process efficiency.

Work Environment

This role involves daily exposure to potentially harmful content, including but not limited to bullying, hate speech, child safety issues, depictions of harm to self/others, and harm to animals. TikTok acknowledges the psychologically demanding nature of this work and provides comprehensive, evidence-based programs to support physical and mental wellbeing throughout employment.

About TikTok

TikTok is the leading destination for short-form mobile video, with a mission to inspire creativity and bring joy. Headquartered in Los Angeles and Singapore, TikTok has global offices in cities including New York, London, Dublin, Paris, Berlin, Dubai, Jakarta, Seoul, and Tokyo.

Benefits

  • Day-one access to medical, dental, and vision insurance.
  • 401(k) savings plan with company match.
  • Paid parental leave, short-term and long-term disability coverage, and life insurance.
  • Wellbeing benefits and 10 paid holidays per year.
  • 10 paid sick days and 17 days of Paid Personal Time annually (prorated upon hire with increasing accruals by tenure).

Pay

The base salary range for this position is $108,800 – $288,000 annually, with compensation varying based on qualifications, skills, experience, and location. Additional compensation may include discretionary bonuses, incentives, and restricted stock units.

Similar jobs

AI Red Team Analyst

AlignerrUnited States· 1 wk ago
RemoteBusiness Developmentapply on alignerr.com