Senior Machine Learning Engineer, AI Evaluation
SHRM is a member-driven catalyst for creating better workplaces where people and businesses thrive together. As the trusted authority on all things work, SHRM is the foremost expert, researcher, advocate, and thought leader on issues and innovations impacting today’s evolving workplaces. With nearly 340,000 members in 180 countries, SHRM touches the lives of more than 362 million workers and their families globally.
About the role
The Senior Machine Learning Engineer, AI Evaluation builds and operates the measurement and engineering infrastructure supporting the organization's Applied AI Research (AAIR) function, a continuous experimental environment designed to evaluate how artificial intelligence models perform real-world HR and workplace-related tasks against established professional standards.
This role is responsible for designing and maintaining the engineering infrastructure used to conduct rigorous, reproducible AI model evaluations and benchmarks. The Senior Machine Learning Engineer develops the systems that run multiple AI models against structured, domain-specific evaluations; builds scoring and evaluation frameworks; maintains reproducibility across model versions; and creates the data infrastructure necessary to analyze and track model performance over time.
Responsibilities
- Evaluate: Identify and surface ambiguity, inconsistencies, or measurement limitations within proposed evaluation criteria and collaborate with subject matter experts to strengthen evaluation design. Evaluate emerging AI models, tools, technologies, and evaluation methodologies and recommend appropriate applications within the research environment.
- Quality Assurance: Ensure evaluation methodologies align with research-defined validation standards and produce findings that are reproducible, transparent, and defensible. Ensure appropriate quality controls are incorporated throughout data collection, evaluation, scoring, storage, and reporting processes. Ensure AI evaluation systems and workflows comply with organizational requirements for data security, privacy, ownership, access, and responsible AI use.
- Maintenance: Maintain portable evaluation architecture across AI model providers to enable consistent and defensible cross-model comparisons as models and technologies evolve. Maintain complete technical documentation and metadata necessary to reproduce research findings and evaluation results. Maintain appropriate controls to protect proprietary, member, research, and other sensitive data from unauthorized access or use.
- Develop: Develop and maintain a unified, provider-agnostic orchestration layer that enables consistent evaluation across multiple frontier model providers and architectures. Develop monitoring, reporting, and visualization capabilities using Looker, Looker Studio, or comparable tools to provide visibility into experiment status, model performance, and performance drift.
- Artificial Intelligence: Apply knowledge of AI evaluation methodologies, benchmarking techniques, inter-rater reliability, and known limitations of automated and model-as-judge evaluation approaches.
- Goals: Establish and maintain technical standards and engineering practices that support reliable, repeatable, and auditable AI evaluation. Establish processes for tracking changes in model behavior across model versions and over time.
- Build: Build systems and processes that support reproducible experimentation, including model-version pinning, comprehensive run logging, experiment tracking, and drift detection. Build and maintain data structures in BigQuery or comparable platforms that enable research results to be queried, analyzed, reproduced, and audited.
- Teamwork: Collaborate with research leaders, HR subject matter experts, data professionals, engineers, and other internal stakeholders to translate research requirements into scalable technical solutions.
- Design: Design, build, and maintain scalable engineering infrastructure for conducting structured evaluations and experiments across multiple AI and large language model (LLM) families. Design and implement rigorous AI evaluation and scoring frameworks, including rubric-based scoring, model-as-judge methodologies with appropriate safeguards, partial-credit methodologies, and approaches for managing ambiguity.
- Support: Support the design of measurement methodologies when definitive ground truth is unavailable or requires expert interpretation. Support collaboration with university, affiliate, research, and other external partners when appropriate and within established organizational access controls and data-handling requirements.
- Technical Contribution: Contribute technical expertise to the design and continuous improvement of AI research experiments, benchmarks, and evaluation methodologies. Implement and maintain appropriate access controls and data-handling requirements when working with external research or partner organizations.
Requirements
- Communication: Demonstrated commitment to reproducibility, including disciplined use of versioning, documentation, logging, experiment tracking, and drift detection. Ability to effectively communicate complex technical concepts, methodologies, limitations, and findings to technical and non-technical audiences.
- Knowledge: A Ph.D. is not required; demonstrated expertise in AI/ML evaluation engineering, research infrastructure, and production-grade systems is valued. Strong knowledge of machine learning, large language models, generative AI systems, and contemporary AI application architectures.
- Experience: Seven (7) or more years of progressively responsible experience in ML/LLM engineering, applied AI, applied data science, or research infrastructure, including experience developing, implementing, and supporting production-grade systems. Demonstrated hands-on experience developing multi-model LLM applications and infrastructure, including provider-agnostic model access, APIs, prompt engineering, and evaluation frameworks.
- Skills: Ability to translate complex, judgment-based requirements from subject matter experts into technically rigorous and measurable evaluation specifications without oversimplifying the underlying domain expertise. Strong analytical and problem-solving skills with the ability to identify technical, methodological, and data-quality issues and develop appropriate solutions.
- Education: Bachelor's degree in Computer Science, Data Science, Machine Learning, Engineering, or a related quantitative or technical field, or relevant equivalent experience in lieu of degree. Master's degree in Computer Science, Data Science, Machine Learning, Artificial Intelligence, or a related field preferred.
- Proficiency: Advanced proficiency in Python and strong software-engineering fundamentals, including the ability to develop reliable, maintainable, production-quality code. Ability to balance technical rigor, research requirements, scalability, and practical implementation considerations. Ability to effectively leverage AI tools and technologies to streamline workflows, enhance productivity, and improve overall work quality.
Schedule
Hybrid Schedule (3 Days In-Office/2 Days Remote). This position follows a hybrid work schedule, with Tuesday through Thursday in office and Monday and Friday remote. Employees must be available during standard business hours, with core hours beginning between 8:00–9:00 a.m. and concluding between 5:00–6:00 p.m. local time. Occasional travel: 0 – 10%.
Pay
$100,000 to $130,000 per year
Physical Requirements
- Prolonged periods of sitting at a desk and working on a computer.
- Frequent use of hands and fingers for typing, handling documents, and using office equipment.
- Occasional standing, walking, bending, and reaching.
- Ability to lift and carry up to 30 pounds as needed.
- Clear verbal and written communication skills for effective interaction with colleagues and stakeholders.