Machine Learning Backend Engineer Graduate (AML MLDev) - 2027 Start
About the role
Data AML is ByteDance's Machine Learning mid-platform, providing training and inference systems for recommendation and advertising for businesses such as Douyin, Jinri Toutiao, and Xigua Video. It delivers powerful Machine Learning computing power to internal business units and conducts research on general and innovative algorithms for business-specific challenges.
Successful candidates must commit to an onboarding date by the end of the year. Please state your availability and graduation date clearly in your resume.
Responsibilities
- Optimize the computational performance of ByteDance's recommendation mid-platform models, conduct in-depth tuning for inference and training bottlenecks in business scenarios, and improve computing utilization.
- Lead the design and development of high-performance kernel libraries, covering general-purpose and business-customized kernels, including general computation and communication parallelism, to ensure ultimate kernel performance.
- Develop model compilation optimization technology, focusing on graph optimization, kernel fusion, computation scheduling, and code generation to build and improve the mid-platform model compilation system.
- Collaborate with business and algorithm teams to identify performance issues, provide full-stack performance analysis, bottleneck diagnosis, and optimization solutions.
- Consolidate general-purpose performance optimization components, toolchains, and platform capabilities to empower multiple internal business units.
Requirements
- Completing or recently completed a Bachelor's or Master's degree in Software Development, Computer Science, Computer Engineering, or a related technical discipline.
- Familiarity with mainstream model compilation stacks (e.g., TVM, MLIR, XLA) and relevant development or optimization experience.
- Proficiency in C/C++ development, with knowledge of assembly, CPU/GPU architecture, and cache mechanisms, and practical experience in high-performance kernel development.
- Understanding of the underlying principles of deep learning frameworks (e.g., TensorFlow, PyTorch, OneFlow), including computational graph structure, inference/training execution, and experience in model graph or compilation optimization.
- Strong independent thinking, problem decomposition, performance troubleshooting, and practical optimization skills to overcome complex performance bottlenecks.
Preferred Qualifications
- Experience in joint hardware and software design, or participation in heterogeneous computing projects.
- Contributions to or development of open-source deep learning kernel libraries, compilers, or inference engines.
- In-depth research experience on the underlying architecture and mechanisms of at least one machine learning framework (e.g., TensorFlow, PyTorch, MxNet, or self-developed frameworks).
Pay
The base salary range for this position is $128,000 – $256,000 annually. Compensation may vary based on qualifications, skills, experience, and location. Eligible roles may include additional discretionary bonuses, incentives, and restricted stock units.
Benefits
- Day-one access to medical, dental, and vision insurance.
- 401(k) savings plan with company match.
- Paid parental leave, short-term and long-term disability coverage, and life insurance.
- Wellbeing benefits.
- 10 paid holidays per year, 10 paid sick days per year, and 17 days of Paid Personal Time (prorated upon hire with increasing accruals by tenure).