Research Scientist - ByteBrain - Global Frontier Tech Recruitment Program - 2027 Start (PhD)
About the Role
This role sits at the intersection of Operations Research, AIOps, and AI for Infrastructure. You will solve challenging optimization problems behind large-scale AI datacenters while pioneering next-generation AI-powered decision-making systems, where LLMs and optimization algorithms work together to improve efficiency, resource utilization, and operational intelligence across ByteDance's global infrastructure.
Responsibilities
- Design and develop AI, machine learning, and optimization algorithms to improve the efficiency, reliability, and performance of large-scale infrastructure systems and AI supply chains (AIOps, operations research, software engineering, AgentOps, system optimization).
- Drive the deployment, scaling, and continuous improvement of algorithms in production environments, supporting large-scale services.
- Identify optimization opportunities and emerging challenges from real-world infrastructure scenarios, translating them into impactful research and engineering solutions.
- Conduct cutting-edge research and publish high-quality papers in top-tier conferences and journals.
Key research areas include:
- Network and Observability: Intelligent fault localization and root cause analysis for large-scale AI clusters, combined with intelligent tuning of time-series databases to improve cluster stability.
- Storage Systems: Develop serverless high-performance elastic file systems and storage acceleration architectures for AI scenarios, explore hardware-software co-optimization for DPU, and overcome AI storage performance bottlenecks.
- Data Center Power Scheduling: Research GPU/CPU/MEM heterogeneous collaborative scheduling technologies, build a heterogeneous power orchestration system for AI agents, and address scheduling challenges including heterogeneous workloads and state dependencies.
- Vector Retrieval: Optimize core vector retrieval technologies for LLM-powered applications, building a cloud-native distributed vector index engine to meet ultra-large-scale vector retrieval demands with low latency and low cost.
- Intelligence and Agent Architecture: Explore automatic infrastructure optimization based on AI Agent workflows, build a self-evolvable business agent framework, and enable full-stack intelligent optimization through AI for Infrastructure.
This work aims to build a next-generation AI-native infrastructure to support LLM and AI agent deployment, improve resource utilization, reduce costs, support elastic scaling, and drive AI infrastructure evolution.
Qualifications
Minimum Qualifications:
- Completing or recently completed a PhD in Computer Science or a related discipline.
- Proven research track record with multiple publications in top-tier conferences or journals related to AI, machine learning, operations research, systems, or related fields.
- Deep expertise in AI, machine learning, and/or operations research, with hands-on experience in large-scale data analysis and algorithm development.
- Strong coding, implementation, and problem-solving skills, with the ability to bridge research and production systems.
- Excellent communication and cross-functional collaboration skills.
Preferred Qualifications:
- Industry experience applying AI and optimization techniques to real-world infrastructure challenges, such as:
- AI data center supply chain optimization
- AI Ops and intelligent operations
- Software engineering productivity optimization
- Operations research and resource scheduling
- System tuning and performance optimization
- Large-scale infrastructure management and automation
Pay
The base salary range for this position is $218,400 - $387,600 annually. Compensation may vary based on qualifications, skills, experience, and location. This role may also be eligible for additional discretionary bonuses, incentives, and restricted stock units.
Benefits
- Day-one access to medical, dental, and vision insurance.
- 401(k) savings plan with company match.
- Paid parental leave, short-term and long-term disability coverage, and life insurance.
- Wellbeing benefits.
- 10 paid holidays per year, 10 paid sick days per year, and 17 days of Paid Personal Time (prorated upon hire with increasing accruals by tenure).