Jobs · Engineering · California

Agent Evaluation & Evolution Machine Learning Engineer Graduate (AML-Ark-US) - 2027 Start

ByteDance · San Jose, CA · 3 wk ago
Engineering$128k–$256k/yrFull-time

About Us

Founded in 2012, ByteDance's mission is to inspire creativity and enrich life. With a suite of products including TikTok, Lemon8, CapCut, Pico, Toutiao, Douyin, and Xigua, ByteDance makes it easier and more fun for people to connect with, consume, and create content. The Applied Machine Learning Ark team combines system engineering and machine learning to develop and operate Large Language Model (LLM) service platforms, offering businesses Model-as-a-Service (MaaS) solutions globally outside mainland China.

Responsibilities

  • Design evaluation systems for LLM-based agents, covering task success, tool use, reasoning quality, and reliability.
  • Build benchmarks and automated judging pipelines, combining rule-based checks, model-based judging, and human review.
  • Analyze agent execution traces and user feedback to identify failure patterns and turn them into concrete system improvements.
  • Support the closed loop from experience to capability, collaborating with research, platform, and product teams to bring methods into production.
  • Develop full-stack, end-to-end solutions spanning text and multimodal LLM algorithms, training/fine-tuning/inference frameworks, prompt engineering, model alignment, and intelligent agent systems.
  • Operate large-scale log analytics pipelines processing invocation logs from text models, multimodal models, and agent systems to extract usage patterns, quality signals, and actionable insights for model improvement and system optimization.

Qualifications

Minimum Qualifications

  • Completing or recently completed a Bachelor's/Master's degree in Computer Science, Artificial Intelligence, Machine Learning, Data Science, or a related field.
  • Solid foundation in machine learning and deep learning.
  • Hands-on experience with LLM-based systems (e.g., agents, tool calling, retrieval, multi-agent systems) through research, internships, or projects.
  • Strong Python skills and experience with a mainstream ML or agent evaluation framework.
  • Demonstrated research or engineering ability through publications, substantial projects, internships, or open-source work.

Preferred Qualifications

  • Publications at top-tier ML/NLP venues (e.g., NeurIPS, ICML, ICLR, ACL), especially in agent learning, self-improving/self-evolving/RSI, or agent evaluation.
  • Experience with evaluation methodology: metric design, model-based judging, or annotation and statistical analysis.
  • Familiarity with LLM post-training, reasoning and planning methods, or continual learning.
  • Experience with feedback-driven optimization loops or large-scale log and trace analysis.

Pay

The base salary range for this position is $128,000 - $256,000 annually. Compensation may vary based on qualifications, skills, experience, and location. This role may also be eligible for additional discretionary bonuses, incentives, and restricted stock units.

Benefits

  • Day one access to medical, dental, and vision insurance.
  • 401(k) savings plan with company match.
  • Paid parental leave.
  • Short-term and long-term disability coverage, and life insurance.
  • Wellbeing benefits.
  • 10 paid holidays per year.
  • 10 paid sick days per year.
  • 17 days of Paid Personal Time (prorated upon hire with increasing accruals by tenure).

Similar jobs