Jobs · Business Development · Washington

Agent Evaluation & Evolution Researcher Graduate (AML-Ark-US) - 2027 Start (PhD)

ByteDance · Seattle, WA · 2 wk ago
Business Development$154k–$301k/yrFull-time

We are the Applied Machine Learning Ark team. We combine system engineering and machine learning to develop and operate Large Language Model (LLM) service platforms that offer businesses Model-as-a-Service (MaaS) solutions, serving both large model providers and downstream users. The US team drives the design, development, and operation of MaaS solutions across the US and international markets outside mainland China. We build full-stack, end-to-end solutions spanning text and multimodal LLM algorithms, training/fine-tuning/inference frameworks, prompt engineering, model alignment, and intelligent agent systems. Beyond model serving, we operate large-scale log analytics pipelines that process massive volumes of invocation logs from text models, multimodal models, and agent systems—extracting usage patterns, quality signals, and actionable insights to inform model improvement, system optimization, and product decisions through continuous, data-driven feedback loops.

Responsibilities

  • Design evaluation systems for LLM-based agents, covering task success, tool use, reasoning quality, and reliability.
  • Build benchmarks and automated judging pipelines, combining rule-based checks, model-based judging, and human review.
  • Analyze agent execution traces and user feedback to identify failure patterns and turn them into concrete system improvements.
  • Support the closed loop from experience to capability, and work with research, platform, and product teams to bring methods into production.

Requirements

  • Individuals who are completing or have recently completed a PhD in Computer Science, Artificial Intelligence, Machine Learning, Data Science, or a related field.
  • Solid foundation in machine learning and deep learning.
  • Hands-on experience with LLM-based systems (e.g., agents, tool calling, retrieval, multi-agent systems) through research, internships, or projects.
  • Strong Python skills and experience with a mainstream ML or agent evaluation framework.
  • Demonstrated research or engineering ability through publications, substantial projects, internships, or open-source work.

Preferred Qualifications

  • Publications at top-tier ML/NLP venues (e.g., NeurIPS, ICML, ICLR, ACL), especially in agent learning, self-improving/self-evolving/RSI, or agent evaluation.
  • Experience with evaluation methodology: metric design, model-based judging, or annotation and statistical analysis.
  • Familiarity with LLM post-training, reasoning and planning methods, or continual learning.
  • Experience with feedback-driven optimization loops, or with large-scale log and trace analysis.

Pay

The base salary range for this position is $153,900 - $300,960 annually. Compensation may vary outside of this range depending on qualifications, skills, competencies, experience, and location. Base pay is one part of the total package, which may include additional discretionary bonuses/incentives and restricted stock units.

Benefits

  • Day-one access to medical, dental, and vision insurance.
  • 401(k) savings plan with company match.
  • Paid parental leave, short-term and long-term disability coverage, and life insurance.
  • Wellbeing benefits.
  • 10 paid holidays per year, 10 paid sick days per year, and 17 days of Paid Personal Time (prorated upon hire with increasing accruals by tenure).

About Us

Founded in 2012, ByteDance's mission is to inspire creativity and enrich life. With a suite of products including TikTok, Lemon8, CapCut, and Pico, as well as platforms specific to the China market like Toutiao, Douyin, and Xigua, ByteDance makes it easier and more fun for people to connect with, consume, and create content.

At ByteDance, we strive to do great things with great people. We lead with curiosity, humility, and a desire to make an impact in a rapidly growing tech company. By fostering an "Always Day 1" mindset, we achieve meaningful breakthroughs for ourselves, our company, and our users.

Similar jobs