Jobs · Engineering · California

Senior Machine Learning Engineer, Agent Eval Platform

ServiceNow · Santa Clara, CA · 2 days ago
On-siteEngineeringFull-time

About the role

The Role
Moveworks' AI agents don't just generate text — they act. They plan, call tools, and change real state in enterprise systems on behalf of 5.5 million employees. That makes the central problem of our team an unusually hard measurement problem: how do you score what an agent did — across a multi-step trajectory through a world it changed — precisely enough that the score can teach it to do better? That signal is what this role owns.

Responsibilities

  • Build the judgement layer of our agent evaluation platform: the rubrics, the judges, the calibration against human labels, the methodology that makes a score mean something.
  • Score that reports its own confidence, so uncertain judgements route to a human instead of quietly becoming training data.
  • Guard against correlated blind spots: our user simulator and our judge are both LLMs, and they can be wrong in the same direction.
  • Hold the guardrail that keeps this honest: step-level scores train and diagnose; end-state outcomes are what we hold the agent to.
  • Turn a calibrated trajectory judge into a process reward model — a dense, step-level signal for what a good agent trajectory looks like — is the unlock.
  • Optimize the agent itself: prompts, tool selection, planner behavior, retrieval, routing — tuned against simulation rather than against production traffic.
  • Create the substrate a future RL effort runs on: versioned scenarios, a repeatable simulated world, and a reward signal calibrated to human judgement.

Requirements

To be successful in this role you have:

  • 5+ years in applied ML, data science, or ML-adjacent engineering, with a track record of work that shipped and got used.
  • Experience turning subjective human judgement into a measurement that holds up — one that other people, and ideally other models, can act on. This is the core of the job.
  • Strong applied ML fundamentals, and comfort treating LLMs as a component you evaluate, prompt, and fine-tune rather than one you pretrain.
  • Python, and the discipline to ship production-grade code rather than notebooks.
  • Ability to think and communicate clearly about complex problems — a large part of this job is convincing engineers that a number means what you say it means, and being right.
  • A high degree of ownership and a bias toward shipping at startup pace.
  • Comfort with ambiguity, and the judgement to know when a measurement is good enough to act on.
  • Experience in at least 3 of these areas: LLM-as-judge or automated evaluation design, and calibrating it against human judgement; Human annotation programs: rubric authoring, label quality, and annotator throughput as a real constraint; Search ranking, recsys, or online experimentation evaluation — golden-set staleness, offline/online divergence, side-by-side rater agreement; Fine-tuning and evaluating small models: SFT, preference tuning, distillation; Reward modeling, RLHF/RLAIF, or process reward models; Agent trajectory analysis and step-level fault attribution; Prompt engineering as an engineering discipline — versioned, tested, and measured, not tuned by vibes.

Qualifications

Experience In At Least 3 Of These LLM-as-judge or automated evaluation design, and calibrating it against human judgement
Human annotation programs: rubric authoring, label quality, and annotator throughput as a real constraint
Search ranking, recsys, or online experimentation evaluation — golden-set staleness, offline/online divergence, side-by-side rater agreement
Fine-tuning and evaluating small models: SFT, preference tuning, distillation
Reward modeling, RLHF/RLAIF, or process reward models
Agent trajectory analysis and step-level fault attribution
Prompt engineering as an engineering discipline — versioned, tested, and measured, not tuned by vibes

Skills

Strong applied ML fundamentals, Python, ability to think and communicate clearly about complex problems, comfort with ambiguity, high degree of ownership, bias toward shipping at startup pace, experience in at least 3 of these areas: LLM-as-judge or automated evaluation design, and calibrating it against human judgement; Human annotation programs: rubric authoring, label quality, and annotator throughput as a real constraint; Search ranking, recsys, or online experimentation evaluation — golden-set staleness, offline/online divergence, side-by-side rater agreement; Fine-tuning and evaluating small models: SFT, preference tuning, distillation; Reward modeling, RLHF/RLAIF, or process reward models; Agent trajectory analysis and step-level fault attribution; Prompt engineering as an engineering discipline — versioned, tested, and measured, not tuned by vibes

Benefits

Equal Opportunity Employer
ServiceNow is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, creed, religion, sex, sexual orientation, national origin or nationality, ancestry, age, disability, gender identity or expression, marital status, veteran status, or any other category protected by law.

Pay

Commensurate with experience

Schedule

TBD

Similar jobs