Staff Software Engineer, Inference Infrastructure
About Us
Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products designed to solve real-world business problems. We train and deploy frontier models for enterprises building AI systems. Our work is instrumental to the widespread adoption of AI, and we are looking for individuals who want to be part of that mission.
We obsess over what we build. Each team member contributes to increasing the capabilities of our models and the value they drive for our customers. Cohere is a global team of researchers, engineers, designers, and more, passionate about their craft. Headquartered in Toronto, we have key offices in London, New York City, San Francisco, Montreal, Paris, Berlin, and Seoul.
About the Role
We are looking for Members of Technical Staff to join the Model Serving team at Cohere. This team is responsible for developing, deploying, and operating the AI platform that delivers Cohere's large language models through easy-to-use API endpoints.
In this role, you will collaborate with multiple teams to deploy optimized NLP models to production in low-latency, high-throughput, and high-availability environments. You will also interface with customers to create customized deployments tailored to their specific needs.
Responsibilities
- Design, deploy, and operate large-scale, highly available distributed systems using Kubernetes and GPU workloads.
- Develop, support, and troubleshoot Kubernetes-based production environments.
- Manage and optimize compute, storage, and network resources across multi-cloud (GCP, Azure, AWS, OCI) and hybrid/on-prem serving environments.
- Ensure smooth operations and efficient teamwork in complex Linux-based computing environments.
- Collaborate with teams to solve evolving technical challenges in mission-critical systems.
- Interface with customers to create and support customized deployments.
Requirements
- 5+ years of engineering experience running production infrastructure at a large scale.
- Experience designing large, highly available distributed systems with Kubernetes and GPU workloads.
- Hands-on experience with Kubernetes development, production coding, and support.
- Experience with GCP, Azure, AWS, OCI, or multi-cloud/on-prem/hybrid serving environments.
- Proficiency in designing, deploying, supporting, and troubleshooting complex Linux-based systems.
- Strong skills in compute, storage, network resource, and cost management.
- Familiarity with computational characteristics of accelerators (GPUs, TPUs, or custom accelerators) and their impact on inference latency and throughput.
- Strong understanding or working experience with distributed systems.
- Experience in Golang, C++, or other high-performance, scalable server languages.
- Excellent collaboration and troubleshooting skills.
- Grit and adaptability to tackle complex, evolving technical challenges.
Benefits
- A weekly lunch stipend of $75/£75 or equivalent in your local currency.
- Full health and dental benefits, including a separate budget for mental health.
- RRSP matching, 401K, or Pension Scheme contributions.
- 100% Parental Leave top-up for up to 6 months for either parent.
- Annual enrichment benefits for arts & culture, fitness/wellness, quality time, and workspace improvements.
- Education and learning stipend for conferences, courses, and coaching.
- 6 weeks of paid vacation (30 working days).
- Budget for traveling to other offices if remote, plus an annual company offsite.
- $500 home office stipend to set up your workspace.
Schedule
This is a full-time position.
Work Arrangements
- For office-based employees: Daily lunch program, plenty of snacks, and regular community/social events.
- For remote employees: Co-working benefit to work alongside others in your city.
- Offices available in Toronto, London, New York City, San Francisco, Montreal, Paris, Berlin, and Seoul, with more locations opening soon.