RL / Post-Training Researcher
About the role
We're partnering with a well-funded AI Fintech Lab tackling some of the most complex real-world problems at the intersection of machine learning, reasoning, and large-scale systems. The team is looking for a Research Scientist focused on reinforcement learning and post-training to help advance model capabilities beyond foundation-model baselines.
You'll work on problems spanning RLHF/RLAIF, preference optimization, reward modeling, synthetic data generation, evaluation frameworks, reasoning enhancement, and large-scale post-training pipelines. The work sits at the boundary of research and production, with a direct path from experimentation to deployment.
This is an opportunity to collaborate with a small group of high-caliber researchers and engineers building novel training systems, developing new approaches to model alignment and optimization, and pushing the limits of what modern agentic AI systems can achieve in real-world environments.
Ideal Experience
- Frontier AI research and engineering
- Reinforcement learning, RLHF, RLAIF, DPO, PPO, GRPO, or reward modeling
- LLM post-training and model alignment
- Evaluation frameworks and model benchmarking
- Synthetic data generation and data flywheel development
- Deep learning and large-scale distributed training
The team is particularly interested in candidates from frontier AI labs, top-tier research organizations, leading technology companies, or high-performing AI startups.
Qualifications
- Degree in CS or a related field from a top-tier institution
Pay
Compensation will be highly competitive to frontier labs, including cash base, discretionary bonus, and equity.