Jobs · Information Technology · Washington

Senior Machine Learning Infrastructure Engineer, Recommendations and Search

TikTok USDS Joint Venture · Seattle, WA · 1 wk ago
On-siteInformation Technology$178k–$342k/yrFull-time

About the Team

We are a group of applied machine learning engineers focusing on TikTok recommendations and search, powering multiple product areas such as For-You-Page (FYP), Live Streaming, Global E-commerce, Local Services, and more. We develop innovative algorithms and techniques to improve user engagement and satisfaction, converting creative ideas into business-impacting solutions. We are passionate about pushing the envelope of State-of-the-Art (SOTA) large-scale machine learning to solve real-world problems.

Responsibilities

  • Technical Execution & System Ownership: Contribute to the technical roadmap and hands-on implementation of our large-scale (in billions of parameters) distributed real-time ML training and inferencing platforms that power the TikTok recommendation and search engines, Short Form Video (SFV) ecosystem. Drive direct business impact on Live, Global E-Commerce, Local Services, and other business domains.
  • Large-Scale Parallelism Architecture: Design and scale multi-node distributed training systems, implementing advanced 3D parallelism strategies (Data, Tensor, Pipeline) to maximize compute efficiency and model scalability. Participate in architecture, scale-testing, and maintenance of massive distributed computing foundations, GPU cluster configurations, and orchestration pipelines to achieve high hardware utilization and cluster efficiency under established SLI/SLO frameworks.
  • Production Inference & Serving: Build and scale low-latency, high-throughput model serving infrastructure, optimizing inference pipelines and leveraging low-level execution paths to handle massive live traffic under strict boundary isolation.
  • Algorithm-Infra Co-Design: Partner closely with Applied ML Research teams to co-design and pioneer next-generation Generative Recommendation systems, abstracting general-purpose components to support advanced generative paradigms in production.
  • Resiliency & Fault Tolerance: Build robust, automated fault-detection systems and asynchronous checkpointing mechanisms to gracefully handle hardware drops, silent data corruption (SDC), or network-switch failures in multi-thousand GPU clusters.
  • Cross-Functional Collaboration: Partner with Applied ML researchers and data platform teams to engineer high-throughput, secure multi-modal data processing and storage engines that prevent compliance friction.
  • Security & Compliance Hardening: Implement and enforce strict encryption-at-rest/in-transit controls, access control lists (ACLs), and secure tenant isolation protocols across the entire compute stack to meet compliance objectives.

Requirements

Minimum Qualifications

  • Bachelor’s or Master's degree in Computer Science, Computer Engineering, or a related technical discipline.
  • 4+ years of professional software engineering experience with deep expertise in Python, C++/Java.
  • 2+ years of direct experience building and maintaining machine learning infrastructure at enterprise scale.
  • Deep technical familiarity with the internals of core ML frameworks (PyTorch / TensorFlow) and a strong understanding of low-level GPU memory management, CUDA interactions, and networking topologies (InfiniBand/RoCE).
  • Solid understanding of production-grade LLM training and inference tools, with hands-on profiling skills to eliminate I/O, compute, or network bottlenecks.
  • Strong system-level troubleshooting and debugging skills, with experience profiling and eliminating I/O, compute, or network bottlenecks.

Preferred Qualifications

  • Experience optimizing high-performance training loops to maximize Model Flops Utilization (MFU) through advanced communication-computation overlap and zero-bubble pipeline scheduling across large-scale distributed clusters using industry-standard frameworks (e.g., Megatron, DeepSpeed).
  • Proven track record of scaling LLM training or inference workloads across hundreds of GPUs, with deep familiarity in advanced serving techniques such as KV Cache management, Prefill-Decoding (PD) separation, and model quantization.
  • Experience working in highly regulated industries, sovereign cloud environments, or dealing with federal data security compliance frameworks.
  • Strong background in low-level kernel development and graph compilation, with experience in CUDA, Triton, Cutlass, TensorRT, or Triton Inference Server.
  • Active interest in keeping up with the latest industry breakthroughs in MLOps, MoE (Mixture of Experts) routing infrastructure, and specialized hardware optimization.
  • Strong communication skills, with the ability to collaborate effectively across team boundaries, draft clear technical design docs, and mentor team members.

About USDS

TikTok USDS Joint Venture LLC is dedicated to the safety and security of millions of Americans who create, discover, and connect with what they love on the apps we operate. The Joint Venture complies with the Executive Order signed by President Trump on September 25, 2025, and operates under a comprehensive data privacy and cybersecurity program with defined safeguards to protect national security and secure U.S. user data, apps, and the algorithm. We safeguard the U.S. content ecosystem, holding decision-making authority for trust and safety policies and moderation.

Why Join Us

Inspiring creativity is at the core of TikTok's mission. Our innovative product helps people authentically express themselves, discover, and connect. We strive to do great things with great people, leading with curiosity, humility, and a desire to make an impact. Every challenge is an opportunity to learn and innovate as one team. We embrace an "Always Day 1" mindset, achieving meaningful breakthroughs for ourselves, our company, and our users.

Benefits

  • Day one access to medical, dental, and vision insurance.
  • 401(k) savings plan with company match.
  • Paid parental leave, short-term and long-term disability coverage, and life insurance.
  • Wellbeing benefits.
  • 10 paid holidays per year.
  • 10 paid sick days per year.
  • 17 days of Paid Personal Time (prorated upon hire with increasing accruals by tenure).
  • Eligibility for additional discretionary bonuses/incentives and restricted stock units.

Pay

The base salary range for this position is $177,688 - $341,734 annually. Compensation may vary based on qualifications, skills, competencies, experience, and location. Base pay is one part of the total compensation package, which may include additional bonuses, incentives, and restricted stock units.

Schedule

The company operates on a fully in-person schedule up to 5 days a week to enable greater speed, alignment, and agility, particularly in areas like real-time decision-making and integrated execution.

Similar jobs