Senior Software Engineer, RL Environments
About the role
You'll own the RL environments frontier labs train on, end to end. Scope the problem with the requester, build the image and the tools inside it, write the graders that score it, ship it into the customer's platform, and keep it healthy once it's running. You sit between Pareto's engineering team and the researchers at the labs we work with, close enough to both that you can tell when a training goal and a buildable spec have drifted apart. Nobody will hand you a finished spec. You'll get a research problem, define what gets built, and stay with it after it lands. In your first year, good looks like environments that ship faster than the last one did, because you invested in the build and release path instead of hand-rolling each delivery. What you build becomes training signal. That's the reason the ownership runs all the way through production.
Responsibilities
- Build industry-leading RL environments and MCP tools that power post-training loops for frontier labs.
- Partner closely with research labs to turn a rough training goal into a spec you can build against, and push back early when the ask won't produce usable signal.
- Investigate failed tasks and jobs and take the lead when a delivery pipeline degrades.
- Automate image builds and release notes, use spec-first design, CI gates, and review harnesses to ensure the next environment costs a fraction of the last one.
- Identify gaps in the platform that fall short of what a lab needs and ensure they get onto the roadmap.
Requirements
- 7+ years of experience building production systems, with the ability to distinguish which problems need careful design and which just need shipping.
- Proficiency in production Python or TypeScript; strong functional-language background is also relevant.
- Experience building and shipping containerized services, with a focus on reproducibility (Docker layering, dependency pinning, consistent image behavior).
- Fluent use of coding agents, with the judgment to review and catch errors in their output.
- Experience owning systems post-launch, including incident response, root cause analysis, and documentation to prevent recurrence.
- Familiarity with cloud infrastructure (containers, managed databases, deploy paths) on AWS or another major cloud platform.
- Based in the US with the ability to travel to the Bay Area as needed.
You Probably Aren't the Right Fit If
- You need requirements locked before you start; specs here evolve with the research.
- You prefer working at arm's length from stakeholders; this role is high-contact by design.
- You want to hand off at launch; this role owns environments in production, including during incidents.
- You only want to build one kind of thing; this role moves between image builds, grader design, and client work in the same week.
Pay
Base salary $245K–$300K, plus equity. Final offer depends on experience and level, and we share level-specific ranges early in the process.
Why Pareto
The environments you build become the training signal for models at Anthropic and GDM. Not adjacent to that work—inside it. Equity is part of the package at every level.
Apply even if you don't match every line above. We care more about how you think than we do about a clean résumé.