Jobs · Quality Assurance

AI Enablement & Governance– AI Quality & Evaluation Lead

Alight Solutions · Illinois, United States · 1 mo ago
RemoteRemoteQuality AssuranceFull-time

At Alight, we believe a company’s success starts with its people. We empower clients to build a healthier and more financially secure workforce by unifying the benefits ecosystem across health, wealth, wellbeing, navigation, and absence management.

Benefits

Alight offers a comprehensive total rewards package, including:

  • Health, dental, and vision coverages starting Day One
  • Wellbeing programs
  • Retirement plans with contribution matching
  • Generous time off and parental leave
  • Continuing education and career growth opportunities

Flexible Working

Alight considers flexible working arrangements wherever possible and has been recognized as a leader in the flexible workspace, named a “Top 100 Company for Remote Jobs” for six consecutive years.

About the Role

The AI Quality & Evaluation Lead enables responsible and scalable AI adoption by defining technical quality standards, evaluation frameworks, and control requirements across the AI lifecycle. This role partners closely with AI Engineering, Data Scientists, QA CoE, and Product teams to embed quality and performance requirements by design, ensuring AI solutions—including RAG systems and Agents—are accurate, grounded, and aligned with enterprise risk and trust expectations.

Responsibilities

  • Quality-by-Design Partnership
    • Partner directly with AI Engineers, Application Developers, and Data Scientists during the design phase to define technical quality acceptance criteria and fit-for-use requirements.
    • Embed quality considerations into model and system architecture, specifically for complex patterns like RAG and autonomous Agents.
    • Define golden truth requirements and evaluation dataset standards; collaborate with Data Science teams to ensure datasets reflect production-level complexity.
    • Set quality and evaluation expectations for third-party AI systems and vendor-supplied models, ensuring consistent governance standards.
  • Technical Evaluation & Metric Engineering
    • Design and maintain a structured evaluation framework to assess AI systems against defined quality bars (e.g., Goodness-of-Fit, Calibration, Stability).
    • Develop automated metrics for Generative AI performance, including Groundedness (Hallucination detection), Faithfulness, Completeness, and other domain-relevant metrics.
    • Define and operationalize fairness and bias evaluation criteria, including demographic parity assessments and disparate impact testing for client-facing AI systems.
    • Calibrate evaluation thresholds and monitoring cadence based on AI risk tier, ensuring proportionate controls.
  • Technical Control & Monitoring
    • Identify and document technical AI governance controls to enable automated compliance with performance and risk obligations.
    • Establish drift and ongoing monitoring requirements, defining statistical triggers for feature and concept drift.
    • Develop clear control statements that articulate expected evidence artifacts (e.g., test results, model cards) for go/no-go decisions.
  • Governance & Evidence Enablement
    • Provide objective, data-driven evaluation outputs to support AI governance reviews and risk classification.
    • Translate governance expectations into clear, testable quality criteria for engineering teams to apply within CI/CD pipelines.
    • Maintain authoritative documentation of AI controls to support audit, regulatory review, and internal assurance activities.

Requirements

  • Technical Depth: 5–8+ years of experience in Data Science, ML Engineering, or AI Quality, with a focus on evaluation and statistical validation.
  • System Design: Practical experience partnering with engineers to design RAG, LLM-based Agents, or traditional ML pipelines.
  • Analytical Skills: Expert-level Python (Pandas, Scikit-learn) and experience with evaluation frameworks (e.g., RAGAS, TruLens, or MLflow).
  • Governance Mindset: Demonstrated ability to translate abstract trust concepts into mathematical metrics and enforceable technical controls.
  • Stakeholder Influence: Ability to drive quality adoption across engineering and product teams without direct authority.
  • Communication: Ability to bridge high-level governance policy and low-level code implementation.
  • Bachelor’s degree in a technical field (e.g., Computer Science, Computer Systems Design) or equivalent professional experience.

Similar jobs