Machine Learning Engineer - Model Evaluation & Experimentation
Weekday AI (YC W21) · United States · 4 wk ago
RemoteRemoteOTHR$60–$90/hrPart-time
Key Responsibilities
- Design realistic machine learning benchmark tasks based on research workflows, including model implementation, experimentation, training, evaluation, and performance analysis
- Translate open-ended research concepts into structured, reproducible evaluation tasks with clearly defined success criteria
- Implement machine learning solutions using Python, execute experiments, and produce reference implementations that demonstrate correct methodology and expected outcomes
- Develop benchmark tasks involving reinforcement learning concepts such as reward functions, policy optimization, training dynamics, and model behavior where applicable
- Evaluate AI-generated solutions by identifying implementation errors, experimental flaws, incorrect reasoning, and unsupported conclusions
- Collaborate with AI researchers and fellow subject matter experts to continuously improve benchmark quality, technical rigor, and evaluation consistency
Required Qualifications
- Master's degree, PhD, or equivalent practical experience in Machine Learning, Computer Science, Artificial Intelligence, Data Science, or another quantitative STEM discipline
- Minimum 1 year of professional experience in machine learning research, research engineering, applied AI, or another research-intensive technical role
- Strong hands-on experience designing, training, evaluating, and optimizing machine learning models through complete experimental workflows
- PRACTICAL experience conducting machine learning experiments, including experiment setup, hyperparameter tuning, execution, validation, and analysis
- Strong understanding of modern Large Language Models (LLMs), their capabilities, limitations, and evaluation methodologies
- Proficiency in Python and Git, with experience working in both script-based and notebook-based development environments
- Familiarity with reinforcement learning concepts—including reward functions, policy optimization, and training behavior—is preferred
- Experience with AI evaluation, benchmark development, AI training, or task authoring is highly desirable
- Excellent analytical thinking, creativity, attention to detail, and the ability to solve complex, open-ended technical problems independently
- Strong written communication skills for documenting experimental methodologies and technical findings
- Able to commit approximately 35 hours per week on a consistent basis
Preferred Qualifications
- Experience developing or evaluating large language models, foundation models, or generative AI systems
- Background in reinforcement learning, deep learning, distributed training, or model optimization
- Familiarity with benchmark design, AI safety evaluations, or research-quality experimentation
- Experience contributing to research publications, open-source machine learning projects, or advanced AI systems