Member of Technical Staff, Training Infra Engineer
About Us
Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products designed to solve real-world business problems. We train and deploy frontier models for enterprises building AI systems. Our work is instrumental to the widespread adoption of AI, and we are looking for individuals passionate about contributing to this mission.
We obsess over what we build. Each team member is responsible for increasing the capabilities of our models and the value they drive for customers. Cohere is a global team of researchers, engineers, designers, and more, headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin, and Seoul.
About the Role
As a Member of Technical Staff, you will contribute to and support model training pipelines, ship state-of-the-art models to production, and bridge the gap between research and production. We have one of the highest ratios of compute to engineers in the world and do not strongly delineate between engineering and research. Everyone contributes to writing production code and supporting research efforts based on individual interest and organizational needs.
We provide all the compute, data, and talent necessary for you to do your best work. This role is remote-friendly with no location restrictions, though we have offices in London, Paris, Toronto, San Francisco, and New York.
Responsibilities
- Design and write high-performance, scalable software for training.
- Improve our training setup from an infrastructure and codebase performance standpoint.
- Craft and implement tools to speed up training cycles and improve the efficacy of our training infrastructure.
- Research, implement, and experiment with ideas on our supercompute and data infrastructure.
- Learn from and collaborate with leading researchers in the field.
Requirements
- Extremely strong software engineering skills.
- Proficiency in Python and related ML frameworks such as JAX, PyTorch, and XLA/MLIR.
- Experience with distributed training infrastructures (Kubernetes, Slurm) and associated frameworks (Ray).
- Experience using large-scale distributed training strategies.
- Hands-on experience training large models at scale and contributing to the tooling and/or setup of the training infrastructure.
- Bonus: Publication at top-tier venues (e.g., NeurIPS, ICML, ICLR, AIStats, MLSys, JMLR, AAAI, Nature, COLING, ACL, EMNLP).
Benefits
- A weekly lunch stipend of $75/£75 or equivalent in your local currency.
- Full health and dental benefits, including a separate budget for mental health.
- RRSP matching, 401K, or Pension Scheme (depending on location).
- 100% parental leave top-up for up to 6 months for either parent.
- Annual enrichment benefits: Arts & culture, fitness/wellness, quality time, and a workspace improvement credit.
- Education & learning stipend for conferences, courses, and coaching.
- 6 weeks of paid vacation (30 working days).
- Budget for traveling to other offices if remote, plus an annual company offsite.
- $500 home office stipend to set up your workspace.
- For those in the office: daily lunch program, plenty of snacks, and regular community and social events.
- For those not near an office: a co-working benefit to work alongside others in your city.