Applied Computer Science Benchmark Specialist
Weekday AI (YC W21) · United States · 1 wk ago
RemoteRemoteOTHR$66–$84/hrContract
Compensation: $66 - $84 per hour
About the role
We are seeking experienced computer science professionals to author and review high-quality academic assessment content for an AI research initiative. In this role, you will develop and validate rigorous multiple-choice questions across a broad range of computer science domains, assess solution quality, and help establish gold-standard benchmarks for evaluating advanced AI systems. You will contribute through one of two primary task types:
- Question Authoring — Develop original, challenging multiple-choice questions within your area of computer science expertise, assess their difficulty, and submit them for review.
- Question Verification — Review existing questions for technical accuracy, clarity, completeness, and rigor. Make necessary edits, assess difficulty, and document the rationale behind your changes.
Responsibilities
- Create original computer science questions that evaluate deep conceptual understanding, technical reasoning, and problem-solving rather than surface-level recall
- Ensure every question is unambiguous, self-contained, technically accurate, and sufficiently specified for a qualified expert to solve
- Classify questions by difficulty:
- Medium: Introductory undergraduate level
- Hard: Advanced undergraduate level
- Expert: Postgraduate level and above
- Provide one correct answer alongside nine plausible but subtly incorrect alternatives designed to distinguish strong technical reasoning from superficial knowledge
- Develop clear, structured solution explanations that demonstrate the reasoning and technical principles required to reach the correct answer
- Provide 1-5 authoritative references per question, drawing from peer-reviewed research, academic publications, university resources, and other reputable technical sources
- For verification assignments, identify issues related to correctness, clarity, completeness, precision, or solvability and clearly explain the reasoning behind any recommended edits
- Apply consistent standards when evaluating questions and solutions to ensure benchmark quality and reproducibility
Requirements
Computer Science Domains:
- Accelerator / GPU Kernel Engineering
- Formal Methods & Automated Reasoning
- Computer Architecture & Accelerators
- Distributed Systems
- DevOps & Site Reliability Engineering
- Data Engineering & Databases
- Cloud Computing & Infrastructure
- Operating Systems & Systems Kernel
- Machine Learning Engineering
- Web & API Development
- Embedded Systems Engineering
- Computer Graphics & Game Development
- Mobile Engineering
Qualifications
- PhD or doctoral candidacy in Computer Science, Electrical Engineering, Computer Engineering, or a closely related discipline
- A Master's degree may be considered for candidates with exceptional expertise in a specialized computer science domain
- Strong command of graduate-level computer science theory, algorithms, systems, software engineering, architecture, and/or machine learning
- Demonstrated depth in one or more of the listed technical domains
- Research publications, substantial industry experience at leading technology organizations, systems engineering experience, or competitive programming experience is a strong plus
- Excellent written English and the ability to communicate complex technical concepts clearly, accurately, and concisely
- Strong attention to detail and the ability to distinguish technically valid solutions from plausible but incorrect approaches
Schedule
- Expected commitment: 10+ hours per week
- Fully remote and asynchronous
- Flexible scheduling based on project requirements
Benefits
- Opportunity to contribute to the development of high-quality benchmarks for evaluating advanced AI systems
- Strong contributors may be considered for additional review, evaluation, or subject-matter expert opportunities
Engagement will be on an independent contractor basis. Payments are made weekly through Stripe or Wise, based on services rendered. H-1B and STEM OPT candidates are not eligible for this opportunity at this time.