Jobs · Engineering

Director, Research - AI Evals

Jobgether · United States · 6 days ago
RemoteRemoteEngineering$258k–$348k/yrFull-time

The Director, Research - AI Evals role is a high-impact leadership opportunity for an experienced research professional passionate about defining and measuring quality in AI-powered products.

About the role

Sit at the intersection of AI, product strategy, and user experience, shaping how next-generation intelligent features are evaluated and improved before reaching millions of users. Establish evaluation frameworks, quality standards, and scalable processes that directly influence product decisions across the organization.

Responsibilities

  • Define and own the evaluation strategy for AI-powered experiences, establishing quality dimensions, success metrics, and decision frameworks.
  • Partner with engineering teams to implement scalable evaluation pipelines, regression testing systems, and reproducible quality measurement processes.
  • Develop dashboards, reporting mechanisms, and executive-ready insights that support product prioritization and go/no-go decisions.
  • Champion consistent AI quality standards across teams and embed evaluation practices into the product development lifecycle.
  • Lead and mentor a small team responsible for executing AI evaluation initiatives in collaboration with internal stakeholders, contractors, and AI-assisted workflows.
  • Influence organizational strategy by translating complex evaluation findings into actionable recommendations for senior leadership.

Requirements

  • 10+ years of experience in product research, applied research, product development, or related disciplines, including at least 2 years of people management experience.
  • Proven hands-on expertise evaluating AI and LLM-powered products, including human evaluation programs and automated evaluation methodologies.
  • Strong understanding of rubric design, benchmark creation, inter-rater reliability, and model-based assessment techniques such as LLM-as-a-judge approaches.
  • Demonstrated ability to identify critical risks, prioritize ambiguous quality questions, and design effective evaluation strategies.
  • Strong stakeholder management and communication skills, with experience influencing executive and cross-functional teams.
  • Familiarity with AI evaluation tooling and infrastructure such as Braintrust, LangSmith, DeepEval, or similar platforms is considered a strong asset.
  • Experience building new functions, practices, or research disciplines from the ground up is highly desirable.
  • Additional background in product design, product management, data science, engineering, or front-end development would be advantageous.

Benefits

  • Competitive compensation package with annual base salary ranging from $258,000 to $348,000 USD, adjusted based on location, experience, and scope.
  • Equity participation program, offering long-term growth and ownership opportunities.
  • Comprehensive health, dental, and vision coverage.
  • Roth retirement plans with company contributions.
  • Generous paid time off, company recharge days, and family-friendly leave policies, including parental and reproductive support programs.
  • Mental health and wellness benefits designed to support overall well-being.
  • Learning and development stipend to encourage continuous professional growth.
  • Work-from-home allowance and cell phone reimbursement to support remote productivity.
  • Flexible and collaborative work environment with opportunities to contribute to cutting-edge AI innovation.

Similar jobs