Site Reliability Engineer, Inference Infrastructure
About Us
Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products designed to solve real-world business problems. We’re training and deploying frontier models for enterprises building AI systems. Our work is instrumental to the widespread adoption of AI, and we’re looking for folks who want to be part of that.
We obsess over what we build. Each team member contributes to increasing the capabilities of our models and the value they drive for customers. Cohere is a global team of researchers, engineers, designers, and more, passionate about their craft. We are co-headquartered in Toronto and San Francisco, with key offices in London, New York City, Montreal, Seoul, Germany, and Paris.
About the Role
We are looking for a Site Reliability Engineer to join the Model Serving team at Cohere. This team is responsible for developing, deploying, and operating the AI platform delivering Cohere's large language models through easy-to-use API endpoints.
In this role, you will work closely with multiple teams to deploy optimized NLP models to production in low-latency, high-throughput, and high-availability environments. You will also interface with customers to create customized deployments that meet their specific needs.
Responsibilities
- Build self-service systems that automate managing, deploying, and operating services, including custom Kubernetes operators that support language model deployments.
- Automate environment observability and resilience to enable all developers to troubleshoot and resolve problems.
- Take steps to ensure we hit defined SLOs, including participation in an on-call rotation.
- Build strong relationships with internal developers and influence the Infrastructure team’s roadmap based on their feedback.
- Develop our team through knowledge sharing and an active review process.
Requirements
- 5+ years of engineering experience running production infrastructure at a large scale.
- Experience designing large, highly available distributed systems with Kubernetes, including GPU workloads on those clusters.
- Experience with Kubernetes development, production coding, and support.
- Experience with GCP, Azure, AWS, OCI, multi-cloud, on-prem, or hybrid serving.
- Experience designing, deploying, supporting, and troubleshooting in complex Linux-based computing environments.
- Experience in compute, storage, network resource, and cost management.
- Excellent collaboration and troubleshooting skills to build mission-critical systems and ensure smooth operations.
- The grit and adaptability to solve complex technical challenges that evolve daily.
- Familiarity with computational characteristics of accelerators (GPUs, TPUs, or custom accelerators), especially how they influence latency and throughput of inference.
- Strong understanding or working experience with distributed systems.
- Experience in Golang, C++, or other languages designed for high-performance scalable servers.
Benefits
- A weekly lunch stipend of $75/£75 or equivalent in your local currency.
- Full health and dental benefits, including a separate budget for mental health.
- RRSP matching, 401K, or Pension Scheme.
- 100% Parental Leave top-up for up to 6 months for either parent.
- Annual enrichment benefits: Arts & culture, fitness/wellness, quality time, and a workspace improvement credit.
- Education & learning stipend for conferences, courses, and coaching.
- 6 weeks of paid vacation (30 working days).
- Budget for traveling to other offices if remote, plus an annual company offsite.
- $500 home office stipend to set up your workspace.
Schedule
Cohere is remote-friendly with offices in Toronto, San Francisco, New York City, London, Paris, Montreal, and more. For those in the office: daily lunch programs, snacks, and regular community events. For those not near an office: a co-working benefit to work alongside others in your city.