Research Scientist Graduate (DPU & AI Infra) - 2027 Start (PhD)
About the Team
The ByteDance DPU (Data Processing Unit) team is building the foundational computing infrastructure for ByteDance and Volcano Engine Public Cloud. Our mission is to advance the architecture, development, and research of next-generation software-hardware technologies across compute, networking, and storage for cloud and AI computing. Our technology stack spans:
- Cloud virtualization & hypervisors
- High-performance user-space network protocols (DPDK, RDMA, etc.)
- High-speed interconnect and virtual switching
- Distributed storage acceleration
- GPU virtualization and scheduling for AI/ML workloads
We work at the intersection of software systems, distributed infrastructure, and custom hardware acceleration, shaping the next wave of cloud-scale computing. We are looking for talented individuals to join our team. As a graduate, you will get opportunities to pursue bold ideas, tackle complex challenges, and unlock limitless growth. Successful candidates must be able to commit to an onboarding date by the end of the year.
Responsibilities
- Design and develop DPU network software with a focus on high performance, low latency, and reliability.
- Collaborate with hardware teams to build software-hardware co-design solutions for networking and storage acceleration.
- Explore AI/ML infrastructure acceleration, leveraging DPUs, GPUs, and custom hardware to optimize distributed training and inference.
- Drive end-to-end performance optimization, from OS kernels and drivers to user-space runtime systems.
- Contribute to architecture design, technical proposals, and long-term research directions.
Qualifications
Minimum Qualifications
- Individuals who are completing or have recently completed a PhD degree in CS or a related discipline.
- Proficiency in C/C++ development and debugging.
- Familiar with Linux systems development experience.
- Solid understanding of compute, network architecture, and operating systems.
- Background in at least one of: software-hardware co-design, distributed systems, high-performance networking, or AI/ML systems.
Preferred Qualifications
- Experience with software-hardware co-design (networking, storage, or distributed compute).
- Hands-on experience with network virtualization (OVS, SR-IOV, eBPF).
- Familiarity with DPDK and high-performance user-space networking.
- Bonus points for hardware acceleration experience, FPGA/ASIC/GPU/CUDA.
- Bonus points for experience with NCCL Collectives along with AI communication patterns and parallelization techniques.
- Proven experience designing and building AI/ML infrastructure related but not limited to inference kv cache system, data preprocessing system.
Pay
The base salary range for this position is $162,000 - $316,800 annually. Compensation may vary outside of this range depending on a number of factors, including a candidate’s qualifications, skills, competencies and experience, and location. Base pay is one part of the Total Package that is provided to compensate and recognize employees for their work; this role may be eligible for additional discretionary bonuses/incentives, and restricted stock units.
Benefits
- Day one access to medical, dental, and vision insurance.
- 401(k) savings plan with company match.
- Paid parental leave.
- Short-term and long-term disability coverage.
- Life insurance.
- Wellbeing benefits.
- 10 paid holidays per year.
- 10 paid sick days per year.
- 17 days of Paid Personal Time (prorated upon hire with increasing accruals by tenure).