Web Research Specialist - LLM Evaluation
Quikr · United States · 2 days ago
RemoteRemoteWriting$60/hrContract
About the role
We are building an evaluation benchmark for frontier AI browsing agents. Your job is to design research problems that a state-of-the-art AI cannot solve, even with full web access and multiple attempts. This is not a subject matter expert role, nor is it a content-writing role. It is investigative research.
Responsibilities
- A natural-language research question with a short, stable, objectively verifiable answer
- Clues, each independently checkable, spanning multiple fact types - including dates, people, places, organisations, works, events, records, and quantities - with specific constraints
- A validation record showing the obvious searches you ran and the results they returned
Requirements
- Demonstrated open-web research ability, including locating primary records and navigating government and institutional databases, archives, registries, and PDF documents
- Precision with sourcing. You cite exact pages, tables, and sections—not just homepages
- Comfort researching unfamiliar subjects from scratch
- Native or near-native written English
- High tolerance for structured documentation. The evidence trail is the majority of the work
- Experience with LLM evaluation, red-teaming, or benchmark construction
- Experience in one or more of the following domains:
- Reference librarianship, archival research, or special collections
- Investigative journalism or professional fact-checking
- OSINT, due diligence, KYC, or investigative research
- Patent, prior-art, or legal-discovery search
- Genealogy and records research
- Competitive quizzing or puzzle-hunt construction
Qualifications
- Nice to have: Experience with LLM evaluation, red-teaming, or benchmark construction
- Familiarity with JSON and structured data delivery formats