GPU/AI Application System Software Engineer Intern (System Technologies and Engineering) - 2027 Summer
About the Team
The GPU/AI System Technology and Engineering Team is committed to developing highly optimized OS and system software to support deep learning and high-performance computing (HPC) workloads in large-scale data centers. We focus on delivering core software components for the next generation of AI and HPC platforms, benchmarks, and fine-tuning performance. Our work spans the entire hardware/software stack, from GPU drivers to deep learning frameworks, to ensure peak performance across all layers. By joining this team, you will work with the best engineers and talents in this industry and have a broad opportunity to get in touch with the latest AI application systems and newly emerged technology in computing, networking and storage. You will gain remarkable GPU architecture, system software development and GPU validation experience in the most advanced hardware infrastructure on a massive scale. We are looking for talented individuals to join us for an internship. Our internship program offers students hands-on experience, industry exposure, and opportunities to apply their knowledge to real-world challenges while building a strong foundation for personal and professional growth. Interns will gain practical experience, explore potential career paths, and participate in social events, learning programs, and development workshops alongside industry professionals.
Responsibilities
- Design and implement performance benchmarks and testing methodologies to evaluate system performance (especially for those factors impacted most closely to OS, OS kernel, Hardware System)
- Develop benchmark tools and performance optimization of AI workloads specifically tailored for large-scale LLM training and inference, as well as High-Performance Computing (HPC)
- Develop Python scripts to automate the testing of various benchmark tools
- Collaborate with internal teams to identify system bottleneck, debug and improve performance issues
Minimum Qualifications
- Currently pursuing a Bachelor's or Master's degree within Computer Engineering in Electrical Engineering, Computer Engineering, Computer Science or related majors
- Deep understanding of Operating System, Linux Kernel, Computer Architecture
- Background with GPU/CPU benchmarking
- Familiar with ML/DL techniques, algorithms and frameworks like TensorFlow or PyTorch
- Exposure to testing automation for various applications
- Hands-on experience with Linux based systems
Preferred Qualifications
- Strong background in one of the following fields: High Performance Computing, ML Hardware Acceleration (e.g., GPU/TPU/RDMA) or ML for Systems, and Distributed Storage
- Experience in AI model development, training, evaluation and deployment on Cloud, Cluster or on-premises
- Experience with parallel programming and at least one communication runtime (MPI, NCCL, UCX, NVSHMEM)
- Linux kernel development experience, such as networking and device drivers etc.
- Exposure to testing automation for various applications
- Experience with complex system-level debugging is invaluable
Benefits
- Day one access to health insurance, life insurance, wellbeing benefits and more
- 10 paid holidays per year
- Paid sick time (56 hours if hired in first half of year, 40 if hired in second half of year)
- Housing allowance for interns who are not working 100% remote
Pay
The hourly rate range for this position in the selected city is $45–$45.