Jobs · Engineering · California

Evals Engineer

Zof AI · San Francisco, CA · Yesterday
EngineeringFull-time

About the role

Zof AI is seeking an Evals Engineer to build the tests that determine whether AI actually works. Verification is the heart of what Zof AI does, so this role sits close to the core of the product: designing eval suites, verification harnesses, and quality gates that confirm AI systems built the right thing, part QA discipline and part domain judgment.

Responsibilities

  • Design and build eval suites for AI products and agent systems.
  • Build verification harnesses that confirm the AI built the right thing.
  • Define quality gates that gate what ships and what does not.
  • Turn domain expertise and customer requirements into testable checks.
  • Hunt failure modes: regressions, hallucinations, and silent errors.
  • Make eval results legible to engineers, product, and customers.
  • Wire evals into CI and the development loop.
  • Raise the standard for what "working" means across the company.

Requirements

  • Experience testing, evaluating, or QA-ing complex software systems.
  • Understanding of how LLM and agent systems fail.
  • Strong analytical rigor and skepticism.
  • Ability to write code to build harnesses and automation.
  • Attention to detail and a high quality bar.
  • Clear written and verbal communication.
  • Comfort operating in a fast-moving environment.
  • High ownership.

Nice to have

  • Experience building LLM evals, benchmarks, or test infrastructure.
  • QA, SDET, or test automation background.
  • Domain expertise in a vertical where correctness matters.
  • Experience with statistical evaluation methods.

Benefits

What we provide in San Francisco:

  • MacBook Pro
  • Premium AI development tools
  • Cursor Ultra Claude Code Ultra OpenAI Codex Max or equivalent advanced AI tooling
  • Access to a high-performance AI product environment
  • Closer collaboration with leadership, engineering, and customers
  • Opportunity to work in the San Francisco AI ecosystem
  • Wellness and productivity support where applicable
  • Competitive startup environment
  • High ownership
  • Direct product impact

Pay

Competitive salary Plus meaningful equity

Schedule

Full-time

Qualifications

Must understand how to measure whether AI systems actually work, beyond demos.

Skills

Must understand how LLM and agent systems fail.

Application

This application is for the Evals Engineer role in San Francisco, CA. Please complete every required field.

Similar jobs

Evaluation Engineer

ElicitOakland, CA· 4 wk ago
RemoteEngineering$140/hrapply on jobs.ashbyhq.com