Product Manager, Evals & Improvement
Meta · Menlo Park, CA · 1 wk ago
Marketing$146k–$204k/yrFull-time
We’re building Metamate, Meta’s internal full-stack platform that lets people direct, review, and help agents improve their work across the company. This role owns how we measure and raise the quality of agentic systems—multi-step trajectories, tool use, and partial credit—in an early-stage, high-leverage space where the right metrics don’t yet exist. You’ll define what “good” means for agents, own the eval verdict that gates model upgrades and major launches, and work shoulder-to-shoulder with engineering, data science, and ML to turn real usage into a continuous improvement loop.
Responsibilities
- Own the end-to-end feedback and improvement system: collection, analysis, prioritization, and resolution.
- Define how we measure agent quality across dimensions (accuracy, helpfulness, reliability, safety).
- Build product mechanisms that capture implicit and explicit user signal at scale.
- Partner with ML and data science to turn feedback into model improvements, eval frameworks, and training data.
- Drive the evals and measurement strategy: what “good” looks like, how we detect regressions, and how we track progress.
- Create transparency and accountability for quality across the ATA Product team.
Requirements
- Strong quantitative skills and experience defining complex metrics.
- Comfort with ambiguous problem spaces where the right metrics do not yet exist.
- 5+ years of relevant industry experience with at least 2 years in Product Management.
- Experience partnering closely with data science and ML engineering teams.
- Track record of driving product improvements through data and experimentation.
- Direct experience building evals for AI products end to end: task set construction, rubric design, instrumentation, and analysis.
- Experience running eval-driven development cycles, with a clear view of where evals are informative and where they mislead.
- Willingness to get into the weeds of eval data and rubrics, with high standards for eval rigor and an obsession with quality.
- Deep understanding of LLM capabilities and failure modes.
Preferred Qualifications
- Experience building annotation, labeling, or crowd-sourcing systems.
- Experience with reinforcement learning from human feedback (RLHF) or similar human-in-the-loop systems.
- Background in evals, trust and safety measurement, or ML quality infrastructure.
- Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements).
- Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews).
- Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies.
Pay
$146,000/year to $204,000/year + bonus + equity + benefits. Individual compensation is determined by skills, qualifications, experience, and location.