Jobs · Engineering · California

Agent Evaluation & Evolution Machine Learning Engineer Intern (AML-Ark-US) - 2027 Summer

ByteDance · San Jose, CA · 2 wk ago
Engineering$45/hrInternship

About Us

Founded in 2012, ByteDance's mission is to inspire creativity and enrich life. With a suite of products including TikTok, Lemon8, CapCut, and Pico, as well as platforms specific to the China market like Toutiao, Douyin, and Xigua, ByteDance makes it easier and more fun for people to connect with, consume, and create content.

The Applied Machine Learning Ark team combines system engineering and machine learning to develop and operate Large Language Model (LLM) service platforms that offer businesses Model-as-a-Service (MaaS) solutions, serving both large model providers and downstream users. The US team drives the design, development, and operation of MaaS solutions across the US and international markets outside mainland China. We build full-stack, end-to-end solutions spanning text and multimodal LLM algorithms, training/fine-tuning/inference frameworks, prompt engineering, model alignment, and intelligent agent systems. Beyond model serving, we operate large-scale log analytics pipelines processing massive volumes of invocation logs to extract usage patterns, quality signals, and actionable insights for model improvement and system optimization.

Responsibilities

  • Design evaluation systems for LLM-based agents, covering task success, tool use, reasoning quality, and reliability.
  • Build benchmarks and automated judging pipelines, combining rule-based checks, model-based judging, and human review.
  • Analyze agent execution traces and user feedback to identify failure patterns and turn them into concrete system improvements.
  • Support the closed loop from experience to capability, working with research, platform, and product teams to bring methods into production.

Qualifications

Minimum Qualifications

  • Currently pursuing a Bachelor's or Master's degree in Computer Science, Artificial Intelligence, Machine Learning, Data Science, or a related field.
  • Solid foundation in machine learning and deep learning.
  • Hands-on experience with LLM-based systems (e.g., agents, tool calling, retrieval, multi-agent systems) through research, internships, or projects.
  • Strong Python skills and experience with a mainstream ML or agent evaluation framework.
  • Demonstrated research or engineering ability through publications, substantial projects, internships, or open-source work.

Preferred Qualifications

  • Publications at top-tier ML/NLP venues (e.g., NeurIPS, ICML, ICLR, ACL), especially in agent learning, self-improving/self-evolving systems, or agent evaluation.
  • Experience with evaluation methodology: metric design, model-based judging, or annotation and statistical analysis.
  • Familiarity with LLM post-training, reasoning and planning methods, or continual learning.
  • Experience with feedback-driven optimization loops or large-scale log and trace analysis.

Benefits

  • Day one access to health insurance, life insurance, and wellbeing benefits.
  • 10 paid holidays per year and paid sick time (56 hours if hired in the first half of the year, 40 if hired in the second half).
  • Eligibility for housing allowance if not working 100% remote.

Pay

The hourly rate range for this internship position is $45–$45.

Similar jobs