Evals Engineer
Zof AI · San Francisco, CA · Yesterday
EngineeringFull-time
About the role
Zof AI is seeking an Evals Engineer to build the tests that determine whether AI actually works. Verification is the heart of what Zof AI does, so this role sits close to the core of the product: designing eval suites, verification harnesses, and quality gates that confirm AI systems built the right thing, part QA discipline and part domain judgment.
Responsibilities
- Design and build eval suites for AI products and agent systems.
- Build verification harnesses that confirm the AI built the right thing.
- Define quality gates that gate what ships and what does not.
- Turn domain expertise and customer requirements into testable checks.
- Hunt failure modes: regressions, hallucinations, and silent errors.
- Make eval results legible to engineers, product, and customers.
- Wire evals into CI and the development loop.
- Raise the standard for what "working" means across the company.
Requirements
- Experience testing, evaluating, or QA-ing complex software systems.
- Understanding of how LLM and agent systems fail.
- Strong analytical rigor and skepticism.
- Ability to write code to build harnesses and automation.
- Attention to detail and a high quality bar.
- Clear written and verbal communication.
- Comfort operating in a fast-moving environment.
- High ownership.
Nice to have
- Experience building LLM evals, benchmarks, or test infrastructure.
- QA, SDET, or test automation background.
- Domain expertise in a vertical where correctness matters.
- Experience with statistical evaluation methods.
Benefits
What we provide in San Francisco:
- MacBook Pro
- Premium AI development tools
- Cursor Ultra Claude Code Ultra OpenAI Codex Max or equivalent advanced AI tooling
- Access to a high-performance AI product environment
- Closer collaboration with leadership, engineering, and customers
- Opportunity to work in the San Francisco AI ecosystem
- Wellness and productivity support where applicable
- Competitive startup environment
- High ownership
- Direct product impact
Pay
Competitive salary Plus meaningful equity
Schedule
Full-time
Qualifications
Must understand how to measure whether AI systems actually work, beyond demos.
Skills
Must understand how LLM and agent systems fail.
Application
This application is for the Evals Engineer role in San Francisco, CA. Please complete every required field.