AI Evaluation Infrastructure Consultant
Landing Point · New York, NY · 3 wk ago
On-siteInformation Technology$120/hrContract
We are a leading financial services company seeking an AI Evaluation Infrastructure Consultant to ensure the quality and compliance of AI systems. This role is crucial for building the evaluation discipline and tooling necessary for AI systems to be accurate, safe, and ready to scale.
Responsibilities
- Develop evaluation harnesses and tooling for AI systems.
- Create golden test sets and scenario libraries for expected behaviors and edge cases.
- Conduct regression testing to identify quality changes.
- Perform hallucination and grounding/faithfulness testing.
- Conduct bias and fairness testing to support fair-lending obligations.
- Implement adversarial and red-team testing.
- Ensure policy-adherence testing against compliance and regulatory requirements.
- Monitor drift detection and ongoing production.
- Manage human and subject-matter-expert evaluation workflows.
- Establish production-readiness gates for AI systems.
- Provide evaluation evidence and reporting for AI governance.
- Define AI vendor acceptance criteria for third-party solutions.
Qualifications
- 7+ years of experience in software quality, data science, machine learning, or related fields.
- Experience designing evaluation methods for LLM or ML systems.
- Strong understanding of testing methodology and statistical rigor.
- Experience with bias and fairness evaluation; familiarity with fair-lending concepts is a plus.
- Ability to define and enforce production-readiness gates and acceptance criteria.
- Experience partnering with risk, compliance, and legal teams.
Pay
$120/hr, depending on experience.