Jobs · Engineering · California

Site Reliability Engineer, Platform Responsibility - USDS

TikTok USDS Joint Venture · San Jose, CA · Yesterday
HybridEngineering$123k–$259k/yrFull-time

Team Intro

The Platform Responsibility engineering team is fast growing and responsible for building machine learning models and systems to identify and defend internet abuse and fraud on our platform. Our mission is to protect billions of users and publishers across the globe every day. We embrace the state-of-the-art machine learning technologies and scale them to detect and improve trust and safety system using the tremendous amount of data generated on the platform. With the continuous efforts from our team, TikTok USDS is able to provide the best user experience and bring joy to everyone in the world.

Responsibilities

  • Manage day-to-day operations of data service, realtime/batch data pipelines, such as SLA/SLO/SLI management, system deployment, performance tuning and troubleshooting
  • Design and deploy AI Agents and LLM-powered automation to streamline incident response, root cause analysis, and proactive system monitoring
  • Create tools and automation to improve system administration and operational efficiency, leveraging AI-assisted development tools to accelerate delivery and code quality
  • Participate in regular on-call rotations as part of a team that provides 24 hour coverage across multiple shifts
  • Engage in and improve the whole lifecycle of services from inception and design, development, capacity planning, and launch reviews, to deployment, operation, and refinement
  • Practice sustainable user support, incident response, and post mortem

Minimum Qualifications

  • Bachelor or above degree in computer science or a related technical discipline
  • At least 1 year of industrial experience
  • Experience integrating AI/LLM APIs into internal workflows or infrastructure tooling
  • Demonstrated independent thinking capabilities and troubleshooting skills
  • Familiar with Unix/Linux system internals, networking, and distributed systems
  • Expertise in monitoring tools (e.g., Prometheus, Grafana, DataDog) and fundamental observability approaches

Preferred Qualifications

  • In-depth knowledge of Unix/Linux systems, networking fundamentals and system performance tuning
  • Familiar with backend systems such as MySQL/Redis/Nginx/Kafka/Kubernetes/Docker and big data technologies such as Hadoop/Spark/Flink/Hive/OLAP/ClickHouse, etc.

About USDS

TikTok USDS Joint Venture LLC is dedicated to the safety and security of millions of Americans who create, discover, and connect with what they love on the apps we operate. The Joint Venture has been established in compliance with the Executive Order signed by President Trump on September 25, 2025. Our foundation is a comprehensive data privacy and cybersecurity program we operate under defined safeguards to protect national security and secure U.S. user data, apps and the algorithm. We safeguard the U.S. content ecosystem, holding decision-making authority for trust and safety policies and moderation. USDS Joint Venture helps ensure Americans can continue to express their creativity, discover new hobbies and interests, and build thriving communities and businesses on a global scale. On-site presence across teams allows the company to operate with greater speed, alignment, and agility — especially in areas like real-time decision-making, team development, and integrated execution. As such, the company is shifting from a hybrid work model to a fully in-person schedule up to 5 days a week.

Why Join Us

Inspiring creativity is at the core of TikTok's mission. Our innovative product is built to help people authentically express themselves, discover and connect – and our global, diverse teams make that possible. Together, we create value for our communities, inspire creativity and bring joy - a mission we work towards every day. We strive to do great things with great people. We lead with curiosity, humility, and a desire to make impact in a rapidly growing tech company. Every challenge is an opportunity to learn and innovate as one team. We're resilient and embrace challenges as they come. By constantly iterating and fostering an "Always Day 1" mindset, we achieve meaningful breakthroughs for ourselves, our company, and our users. When we create and grow together, the possibilities are limitless. Join us.

Pay & Benefits

The base salary range for this position in the selected city is $122,574 - $259,200 annually. Compensation may vary outside of this range depending on a number of factors, including a candidate's qualifications, skills, competencies and experience, and location. Base pay is one part of the Total Package that is provided to compensate and recognize employees for their work, and this role may be eligible for additional discretionary bonuses/incentives, and restricted stock units.

Employees have day one access to medical, dental, and vision insurance, a 401(k) savings plan with company match, paid parental leave, short-term and long-term disability coverage, life insurance, wellbeing benefits, among others. Employees also receive 10 paid holidays per year, 10 paid sick days per year and 17 days of Paid Personal Time (prorated upon hire with increasing accruals by tenure). The Company reserves the right to modify or change these benefits programs at any time, with or without notice.

Similar jobs

Site Reliability Engineer

Tata Consultancy ServicesScottsdale, AZ· 1 wk ago
Engineering$100k–$110k/yrapply on ibegin.tcsapps.com