Jobs · OTHR

Head of Evaluations (Legal AI Benchmarking)

Newcode.ai · New York, NY · 1 wk ago
RemoteRemoteOTHRFull-time

Position Overview

We are seeking a highly analytical professional with a strong statistical background to join our Head of Evaluations. In this role, you will design, implement, and scale the testing frameworks used to evaluate our platform. You will ensure our AI products meet the highest standards of legal reasoning, factual accuracy, and regulatory compliance while maintaining a near-zero hallucination rate.

Key Responsibilities

  • Design Legal Benchmarks for: Contract Drafting, Information Extraction, Legal Research, and Contract Review

  • Build, source and maintain relevant datasets

  • Audit AI Output: Review and score complex AI-generated legal text, contract analyses, and statutory interpretations for accuracy and precision and lay out a strategy.

  • Define Evaluation Metrics: Establish clear criteria for grading model performance, specifically focusing on logical reasoning, citation accuracy, and the model's ability to safely abstain from answering.

  • Collaborate with Engineering: Partner directly with Engineering to translate legal errors into actionable technical feedback for model fine-tuning.

Requirements

  • PhD or Masters in statistics, mathematics, machine learning or equivalent

  • Analytical Skills: Proven ability to break down complex statutory frameworks and case law into structured, logical data points.

  • Tech-Savviness: python, panda, numpy, jupiter notebooks and similar statistical models

Similar jobs

Legal AI Solutions

Dinsmore & Shohl LLPCincinnati, OH· 2 wk ago
apply on jobs.dayforcehcm.com

AI Legal Engineer

Holland & Knight LLPTampa, FL· 3 days ago
Engineering$163k–$245k/yrapply on hklaw.wd1.myworkdayjobs.com