AI Safety Evaluation & Governance Product Manager Intern (TikTok-Platform Responsibility-Feed Safety) - 2027 Summer
About the Team
The Feed Safety-Model & Data Intelligence Team within TikTok Platform Responsibility ensures that AI models meet the highest bar before they make content safety decisions affecting billions of users. Our work spans three layers:
- Standards & Governance — We define and iterate the safety standards that AI systems must follow, translating complex policy intent into structured, machine-interpretable frameworks. This requires deep governance thinking: navigating trade-offs between safety, fairness, user experience, and enforcement consistency.
- AI/ML Solution Design — We partner closely with algorithm teams to improve model accuracy and stability across safety scenarios, tackling challenges unique to this domain — adversarial content, imbalanced distributions, and deep contextual understanding. We evaluate, select, and help shape the right AI approaches (LLMs, prompting strategies, agentic workflows, etc.) for each problem.
- Rigorous Evaluation — We design statistically grounded evaluation frameworks, build high-quality ground truth datasets, and ensure our assessments are valid, reproducible, and actionable — so the platform can confidently ship AI-powered safety systems at scale.
Our work sits at the intersection of AI/ML product development, trust & safety policy, and data-driven quality assurance — ensuring that AI systems can be reliably deployed for high-precision content review and risk governance at scale.
Our internship program offers students hands-on experience, industry exposure, and opportunities to apply their knowledge to real-world challenges while building a foundation for personal and professional growth. Interns will gain practical experience, explore potential career paths, and participate in social events, learning programs, and development workshops alongside industry professionals.
Responsibilities
- Design and execute evaluation plans for AI safety models — define evaluation objectives, select appropriate metrics, and determine what "good" looks like for each use case.
- Build and maintain high-quality ground truth datasets — design data sampling strategies, develop data cleaning pipelines, and ensure labeling consistency and accuracy.
- Analyze model performance using statistical methods (sampling design, confidence intervals, error analysis) to produce actionable insights for algorithm teams and stakeholders.
- Collaborate with algorithm engineers to translate evaluation findings into concrete model improvement directions; participate in prompt design and model configuration iteration.
- Communicate evaluation results and governance standards to cross-functional partners (Policy, Operations, Algorithm); align on definitions and help calibrate quality expectations.
- Continuously improve evaluation processes — identify gaps, propose methodology upgrades, and ensure our evaluation systems scale with model and policy evolution.
Requirements
- Currently pursuing an Undergraduate/Master's in Statistics, Computer Science, Data Science, Public Policy, or closely related quantitative fields.
- Solid grasp of applied statistics — sampling, hypothesis testing, confidence intervals, distribution analysis — and ability to apply these to real measurement problems.
- Foundational understanding of AI/ML concepts (classification, NLP, LLMs, precision/recall); comfortable discussing model behavior with engineers.
- Interest in governance, policy, or content safety; appreciation for the complexity of defining "right" and "wrong" at scale.
- Strong structured thinking — able to decompose ambiguous problems into clear goals and prioritized actions; goal-oriented and hypothesis-driven.
- Ability to communicate in both English and Chinese to collaborate with global and China-based stakeholders.
Preferred Qualifications
- Internship or project experience in Trust & Safety, AI/ML product, model evaluation, or policy-related work.
- Hands-on experience with prompt engineering, LLM-based evaluation, or building evaluation datasets.
- Familiarity with safety-specific challenges: adversarial content, inter-annotator disagreement, or human-in-the-loop systems.
- Coursework or research in AI governance, computational social science, or interdisciplinary areas combining technology and policy.
- Experience with experimental design and statistical modeling beyond introductory level.
Pay
The hourly rate range for this position is $35–$35.
Benefits
- Day one access to health insurance, life insurance, and wellbeing benefits.
- 10 paid holidays per year and paid sick time (56 hours if hired in the first half of the year, 40 if hired in the second half).
- Eligibility for housing allowance (for interns not working 100% remote).