Jobs · Information Technology

Staff Software Engineer, Inference Infrastructure

Cohere · New York, NY · 1 wk ago
RemoteRemoteInformation TechnologyFull-time

About Us

Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products designed to solve real-world business problems. We train and deploy frontier models for enterprises building AI systems. Our work is instrumental to the widespread adoption of AI, and we are looking for individuals who want to be part of that mission.

We obsess over what we build. Each team member contributes to increasing the capabilities of our models and the value they drive for our customers. Cohere is a global team of researchers, engineers, designers, and more, passionate about their craft. Headquartered in Toronto, we have key offices in London, New York City, San Francisco, Montreal, Paris, Berlin, and Seoul.

About the Role

We are looking for Members of Technical Staff to join the Model Serving team at Cohere. This team is responsible for developing, deploying, and operating the AI platform that delivers Cohere's large language models through easy-to-use API endpoints.

In this role, you will collaborate with multiple teams to deploy optimized NLP models to production in low-latency, high-throughput, and high-availability environments. You will also interface with customers to create customized deployments tailored to their specific needs.

Responsibilities

  • Design, deploy, and operate large-scale, highly available distributed systems using Kubernetes and GPU workloads.
  • Develop, support, and troubleshoot production infrastructure in complex Linux-based computing environments.
  • Manage compute, storage, network resources, and costs efficiently.
  • Collaborate with cross-functional teams to ensure smooth operations and mission-critical system reliability.
  • Solve evolving technical challenges with grit and adaptability.
  • Optimize inference latency and throughput, considering the computational characteristics of accelerators (GPUs, TPUs, or custom accelerators).

Requirements

  • 5+ years of engineering experience running production infrastructure at a large scale.
  • Experience designing large, highly available distributed systems with Kubernetes and GPU workloads.
  • Hands-on experience with Kubernetes development, production coding, and support.
  • Experience with GCP, Azure, AWS, OCI, or multi-cloud/on-prem/hybrid serving environments.
  • Strong understanding or working experience with distributed systems.
  • Proficiency in Golang, C++, or other high-performance, scalable server languages.
  • Excellent collaboration and troubleshooting skills.

Benefits

  • A weekly lunch stipend of $75/£75 or equivalent in your local currency.
  • Full health and dental benefits, including a separate budget for mental health.
  • Retirement savings matching (RRSP, 401K, or Pension Scheme).
  • 100% parental leave top-up for up to 6 months for either parent.
  • Annual enrichment benefits for arts & culture, fitness/wellness, quality time, and workspace improvements.
  • Education and learning stipend for conferences, courses, and coaching.
  • 6 weeks of paid vacation (30 working days).
  • Budget for traveling to other offices if remote, plus an annual company offsite.
  • $500 home office stipend to set up your workspace.

Schedule

This is a full-time position.

Work Arrangements

Cohere is remote-friendly with offices in Toronto, London, New York City, San Francisco, Montreal, Paris, Berlin, and Seoul.

  • For office-based employees: daily lunch program, snacks, and regular community/social events.
  • For remote employees: co-working benefit to work alongside others in your city.

Similar jobs