Director, Research - AI Evals
Jobgether · United States · 6 days ago
RemoteRemoteEngineering$258k–$348k/yrFull-time
The Director, Research - AI Evals role is a high-impact leadership opportunity for an experienced research professional passionate about defining and measuring quality in AI-powered products.
About the role
Sit at the intersection of AI, product strategy, and user experience, shaping how next-generation intelligent features are evaluated and improved before reaching millions of users. Establish evaluation frameworks, quality standards, and scalable processes that directly influence product decisions across the organization.
Responsibilities
- Define and own the evaluation strategy for AI-powered experiences, establishing quality dimensions, success metrics, and decision frameworks.
- Partner with engineering teams to implement scalable evaluation pipelines, regression testing systems, and reproducible quality measurement processes.
- Develop dashboards, reporting mechanisms, and executive-ready insights that support product prioritization and go/no-go decisions.
- Champion consistent AI quality standards across teams and embed evaluation practices into the product development lifecycle.
- Lead and mentor a small team responsible for executing AI evaluation initiatives in collaboration with internal stakeholders, contractors, and AI-assisted workflows.
- Influence organizational strategy by translating complex evaluation findings into actionable recommendations for senior leadership.
Requirements
- 10+ years of experience in product research, applied research, product development, or related disciplines, including at least 2 years of people management experience.
- Proven hands-on expertise evaluating AI and LLM-powered products, including human evaluation programs and automated evaluation methodologies.
- Strong understanding of rubric design, benchmark creation, inter-rater reliability, and model-based assessment techniques such as LLM-as-a-judge approaches.
- Demonstrated ability to identify critical risks, prioritize ambiguous quality questions, and design effective evaluation strategies.
- Strong stakeholder management and communication skills, with experience influencing executive and cross-functional teams.
- Familiarity with AI evaluation tooling and infrastructure such as Braintrust, LangSmith, DeepEval, or similar platforms is considered a strong asset.
- Experience building new functions, practices, or research disciplines from the ground up is highly desirable.
- Additional background in product design, product management, data science, engineering, or front-end development would be advantageous.
Benefits
- Competitive compensation package with annual base salary ranging from $258,000 to $348,000 USD, adjusted based on location, experience, and scope.
- Equity participation program, offering long-term growth and ownership opportunities.
- Comprehensive health, dental, and vision coverage.
- Roth retirement plans with company contributions.
- Generous paid time off, company recharge days, and family-friendly leave policies, including parental and reproductive support programs.
- Mental health and wellness benefits designed to support overall well-being.
- Learning and development stipend to encourage continuous professional growth.
- Work-from-home allowance and cell phone reimbursement to support remote productivity.
- Flexible and collaborative work environment with opportunities to contribute to cutting-edge AI innovation.