Research Program Manager - Research Infrastructure
About the role
Research Program Managers at Reflection are high-leverage leaders and operators who embed directly with research and infrastructure teams to accelerate the pace of frontier model development. They own cross-functional programs spanning training infrastructure and cluster reliability across pre-training, mid-training, and post-training workstreams. They drive end-to-end coordination scaling our training stack alongside engineering leads and external partners. They jump into active incidents and escalations to triage, coordinate response, and drive resolution across teams.
Responsibilities
- Own cross-functional programs spanning training infrastructure and cluster reliability across pre-training, mid-training, and post-training workstreams.
- Drive end-to-end coordination scaling our training stack alongside engineering leads and external partners.
- Jump into active incidents and escalations to triage, coordinate response, and drive resolution across teams.
- Champion a culture of blameless post-mortems and continuous learning, turning every incident into a concrete improvement to our systems and processes.
- Partner with infrastructure and research engineering leads to identify bottlenecks, define priorities, and ensure that infrastructure investments are directly tied to research velocity.
- Create visibility into training run health, cluster reliability, and infrastructure performance so that leadership and teams have the context they need to make fast, informed decisions.
- Create lightweight, durable processes for cross-team handoffs, config management, checkpoint workflows, and other coordination-heavy touchpoints that currently rely on ad hoc communication.
- Translate technical complexity into clear status updates and decision frameworks for engineering leadership and executives.
Requirements
7+ years of experience in technical program management, research operations, or infrastructure coordination, ideally in ML/AI or large-scale distributed systems environments. Deep technical knowledge to engage with engineers on topics like distributed training frameworks, GPU cluster architecture, scheduler behavior, networking, and storage systems. Proven ability to operate effectively in high-ambiguity, fast-moving environments. Strong stakeholder management skills across both deeply technical ICs and senior leadership. Excited to build from zero to one.
Qualifications
- Deep technical knowledge to engage with engineers on topics like distributed training frameworks, GPU cluster architecture, scheduler behavior, networking, and storage systems.
- Proven ability to operate effectively in high-ambiguity, fast-moving environments.
- Strong stakeholder management skills across both deeply technical ICs and senior leadership.
- Excited to build from zero to one.
Skills
Motivated by enabling researchers and engineers to build the world's most capable open-weight AI systems.
Benefits
We believe that to make intelligence open and accessible to all, you need to start at the foundation. Joining Reflection means building from the ground up as part of a talent-dense team. You will help define our future as a company, and help define the future of open foundational models.
Pay
Top-tier compensation: Salary and equity structured to recognize and retain our talent globally. Stock options: Everyone who joins and contributes to Reflection's success gets to share in the upside through stock options. Comprehensive medical, dental, vision, and life, with an annual wellness allowance. Lunch and dinner are provided in the office daily. 22 weeks paid parental leave for all new birthing and non-birthing parents, including adoptive and surrogate journeys. Unlimited paid time off in the U.S. and 30 days in the U.K.
Schedule
Unspecified.