Jobs · Engineering · California

Machine Learning System Scheduling Engineer Graduate (Applied Machine Learning) - 2027 Start

ByteDance · San Jose, CA · 2 wk ago
Engineering$128k–$256k/yrFull-time

Volcano Ark is an all-in-one large model service platform launched by Volcano Engine. It is a leading platform in China's large model market by product capability and market share. The platform provides end-to-end services including model inference, evaluation, fine-tuning, AI application development, and a plugin ecosystem. Volcano Ark hosts Doubao and leading industry large models, and supports enterprise AI adoption through stable, secure, and trusted solutions as well as professional algorithm and technical services.

Data AML is ByteDance's machine learning platform team. It provides training and inference systems for recommendation, advertising, computer vision, speech, and NLP scenarios across products such as Douyin, Toutiao, and Xigua Video. The team also supports internal business teams with large-scale machine learning compute, explores general and innovative algorithms for business problems, and offers core machine learning and recommendation system capabilities to external enterprise customers through Volcano Engine.

Responsibilities

  • Design and develop resource scheduling systems for machine learning workloads, supporting Volcano Ark and machine learning platform products.
  • Optimize orchestration and scheduling of heterogeneous compute resources, including GPUs, CPUs, and other accelerators, as well as storage resources such as cloud storage and networking resources such as VPC and RDMA, across multiple data centers and clusters.
  • Support scheduling requirements for offline training, online inference, and other workloads under strict multi-tenant isolation, improving overall resource utilization and efficiency.

Qualifications

Minimum Qualifications

  • Individuals who are completing or have recently completed a Bachelor's or Master's degree in Computer Science or a related discipline.
  • Proficient in one or two programming languages in a Linux environment, such as Go, Java, or Python.
  • Solid foundation in computer science and programming, familiarity with common algorithms and data structures, and good coding habits.
  • Familiar with at least one mainstream machine learning framework, such as TensorFlow, PyTorch, or an internally developed framework.
  • Familiar with Kubernetes architecture and ecosystem, as well as container technologies such as Docker, container, and Kata.
  • Understands distributed system principles and has participated in the design, development, or maintenance of large-scale distributed systems.

Preferred Qualifications

  • Practical experience in large-scale cluster online/offline resource scheduling; source-level understanding of one or more open-source schedulers such as Kubernetes, Volcano, YARN, or Mesos; familiarity with containerization and lightweight virtualization technologies.
  • Deep understanding and practical experience in scheduling topics such as multi-tenant quota governance, preemption, elasticity, fragmentation, tidal scheduling, co-location, and QoS; strong analytical and modeling ability for complex problems; GPU scheduling experience is preferred.
  • Experience in at least one of the following areas: CUDA, RDMA, AI infrastructure, hardware and software co-design, high-performance computing, machine learning hardware architecture such as GPUs, accelerators and networking, ML for systems, or distributed storage.
  • Hands-on experience in cloud-native machine learning systems is preferred.

Pay

The base salary range for this position in the selected city is $128,000 - $256,000 annually. Compensation may vary outside of this range depending on a number of factors, including a candidate’s qualifications, skills, competencies and experience, and location. Base pay is one part of the Total Package that is provided to compensate and recognize employees for their work, and this role may be eligible for additional discretionary bonuses/incentives, and restricted stock units.

Benefits

  • Day one access to medical, dental, and vision insurance.
  • 401(k) savings plan with company match.
  • Paid parental leave, short-term and long-term disability coverage, life insurance, wellbeing benefits.
  • 10 paid holidays per year, 10 paid sick days per year, and 17 days of Paid Personal Time (prorated upon hire with increasing accruals by tenure).

About Us

Founded in 2012, ByteDance's mission is to inspire creativity and enrich life. With a suite of more than a dozen products, including TikTok, Lemon8, CapCut and Pico as well as platforms specific to the China market, including Toutiao, Douyin, and Xigua, ByteDance has made it easier and more fun for people to connect with, consume, and create content.

Why Join ByteDance

Inspiring creativity is at the core of ByteDance's mission. Our innovative products are built to help people authentically express themselves, discover and connect – and our global, diverse teams make that possible. Together, we create value for our communities, inspire creativity and enrich life - a mission we work towards every day.

As ByteDancers, we strive to do great things with great people. We lead with curiosity, humility, and a desire to make impact in a rapidly growing tech company. By constantly iterating and fostering an "Always Day 1" mindset, we achieve meaningful breakthroughs for ourselves, our Company, and our users. When we create and grow together, the possibilities are limitless.

Similar jobs