Software Engineer - AI Agent Evaluation (Remote)
Mindrift · NAMER · Today
RemoteRemoteWriting$50/hrPart-time
What this opportunity involves
- Create realistic developer environments
- Craft tasks and define evaluation criteria
- Write tests to verify agent solutions
- Iterate on tasks and tests based on feedback
What this is NOT
- Data labeling
- Prompt engineering
- Writing code from scratch
What we look for
- 5+ years in software development
- Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis
- Experience writing tests (functional, integration)
- English proficiency: B2+
Why this is hard
- Frontier models are already good at coding
- Creating a task that genuinely challenges the best models is non-trivial
- You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution
- Tasks have many valid solutions
- Writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds
How it works
- Apply
- Pass qualification(s)
- Join a project
- Complete tasks
- Get paid
Compensation
Up to $50/hr equivalent, depending on level and pace.
Tasks are estimated at ~20 hours each; you set your own schedule.