Jobs · OTHR

Research Scientist

OpenRouter · United States · 1 wk ago
RemoteRemoteOTHRFull-time

About the role

The Research Scientist will conduct deep, original research that advances how the world understands, evaluates, and routes large language models. They will work with one of the richest datasets in AI: billions of LLM generations spanning every major model, provider, and use case. The role involves owning and pursuing a research agenda, designing experiments, developing evaluation frameworks, and producing work that shapes how models are compared, selected, and deployed.

What You'll Do

  • Own and pursue a research agenda focused on LLM evaluation, model quality, routing optimization, and AI usage patterns, contributing original insights that advance the field.
  • Design novel evaluation frameworks and benchmarks that go beyond standard leaderboards, using real-world generation data to capture how models actually perform across tasks and contexts.
  • Conduct large-scale empirical studies on LLM behavior: how models compare across providers, how performance changes over time, and how usage patterns reveal strengths and weaknesses.
  • Develop the statistical and mathematical foundations behind our routing systems, building the models and heuristics that power intelligent provider and model selection.
  • Identify opportunities to apply research findings to feed back into OpenRouter's product and platform.
  • Collaborate with external researchers, model providers, and the open-source community to advance shared understanding of LLM capabilities and limitations.
  • Work with product and engineering teams to translate research findings into improvements to OpenRouter's platform, without being constrained to a shipping cadence.

What You Bring

  • Experience & Technical Skills:
    • MS or PhD in a quantitative field (machine learning, statistics, computer science, mathematics, computational linguistics, or similar).
    • Track record of original research, demonstrated by first-author publications, significant open-source contributions, or equivalent impact in industry research.
    • Deep expertise in statistics, experimental design, and causal inference.
    • You can design rigorous studies and reason carefully about validity, bias, and generalizability.
    • Strong programming skills in Python.
    • You can build data pipelines, run large-scale experiments, and prototype models efficiently.
    • Proficiency in SQL for working with large-scale analytical databases (ClickHouse, BigQuery, or similar).
    • Hands-on experience with modern ML/NLP techniques such as LLM evaluation, fine-tuning, embeddings, classification, or reinforcement learning from human feedback.
    • Familiarity with the current LLM landscape: model architectures, provider ecosystems, benchmark suites, and the strengths and limitations of leading models.
  • Mindset & Approach:
    • Deeply curious and self-directed. You identify the most important open questions and pursue them without waiting for direction.
    • Rigorous but pragmatic. You hold yourself to high scientific standards while operating at startup speed.
    • AI-first in your own workflow. You use LLMs, coding agents, and modern AI tools heavily in your research process and have strong opinions about what works.
    • Strong communicator. You can explain complex findings clearly in papers, blog posts, internal memos, and conversations with non-technical stakeholders.
    • Collaborative. You work well with product and engineering teams and can translate research insights into actionable recommendations.

Similar jobs

Research Scientist

The Henry M. Jackson Foundation for the Advancement of Military MedicineCalifornia, United States· 1 wk ago
OTHR$89k–$140k/yrapply on jobs.hjf.org

Research Scientist

Skild AISan Francisco, CA· 3 wk ago
OTHR$100k–$300k/yrapply on grnh.se

Research Scientist

Anduril IndustriesHuntsville, AL· 2 wk ago
OTHR$165k–$218k/yrapply on boards.greenhouse.io

Research Scientist

Anduril IndustriesBroomfield, CO· 2 wk ago
OTHR$126k–$167k/yrapply on boards.greenhouse.io