Jobs · OTHR

Software Engineer - AI Agent Evaluation (Remote)

Mindrift · United States · Yesterday
RemoteRemoteOTHR$50/hrPart-time

About the role

Mindrift connects specialists with project-based AI opportunities for leading tech companies, focusing on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment.

Responsibilities

  • Create challenging tasks and evaluation criteria within realistic simulated environments
  • Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history
  • Craft the prompt, define what "solved" means, and ensure the task is solvable by an AI agent
  • Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient
  • Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust

Requirements

  • 5+ years in software development
  • Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis
  • Experience writing tests (functional, integration)
  • English proficiency - B2+

What this is NOT

  • Data labeling
  • Prompt engineering
  • Writing code from scratch - the agent writes most of the code; you guide and evaluate

Why this is hard

Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non-trivial. You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution. Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.

How It Works

  • Apply → Pass qualification(s) → Join a project → Complete tasks → Get paid
  • Tasks are estimated at :20 hours each; you set your own schedule.

Pay

Up to $50/hr equivalent, depending on level and pace.

Similar jobs