Human Data Quality Engineer (Founding Team)
Prolific · Seattle, WA · 2 days ago
RemoteRemoteEngineeringFull-time
About the role
The Human Data Quality Engineer will play a pivotal role in defining and implementing quality frameworks for complex human data programs. This role requires a strong foundation in Python and SQL, along with a deep understanding of machine learning pipelines and their impact on model performance.
Responsibilities
- Design and implement quality frameworks for human data programs, including evaluation rubrics and launch readiness strategies.
- Work closely with clients to translate their model needs into robust human data and evaluation strategies.
- Engage with product and engineering teams to build scalable quality systems and infrastructure.
- Investigate and address data quality and integrity issues, and develop scalable solutions.
- Develop and maintain dashboards and monitoring tools to track quality and operational performance.
- Lead the quality improvement efforts across the organization, including upskilling junior analysts.
- Shape best practices for quality in emerging AI domains as the team grows.
Requirements
- 5+ years of experience in building quality, evaluation, or annotation systems within AI, machine learning, LLMs, or human data environments.
- Strong Python and SQL skills, with a passion for using data to solve complex quality problems.
- A solid understanding of machine learning pipelines and how human data impacts model performance.
- Strong analytical and statistical thinking, with experience designing scalable quality frameworks.
- The confidence and credibility to interact with stakeholders at frontier labs and act as a partner.
- The ability to leverage your experience and expertise to influence and guide stakeholders at every level, both client side and internally.
- The ability to turn your own data analysis and quality methodology into requirements that product and engineering can build into systems.
- The ability to explain your data analysis and findings clearly to non-technical stakeholders, so they can act on them.
- A proactive, builder's mindset - you enjoy creating new systems, navigating ambiguity, and improving how things work.
Preferred Skills
- Experience with LLM evaluation, RLHF, AI safety, or red teaming.
- Experience translating vague model or evaluation goals into clear annotation specifications.
- Experience working with human annotation programs or human data operations.
- Familiarity with calibration, inter-rater agreement, drift detection, or other evaluation methodologies.