Jobs · Engineering

Member of Technical Staff, Model Efficiency

Cohere · San Francisco, CA · 2 wk ago
RemoteRemoteEngineeringFull-time

About Us

Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products designed to solve real-world business problems. We’re training and deploying frontier models for enterprises building AI systems. Our work is instrumental to the widespread adoption of AI, and we’re looking for folks who want to be part of that.

We obsess over what we build. Each team member contributes to increasing the capabilities of our models and the value they drive for customers. Cohere is a global team of researchers, engineers, designers, and more, passionate about their craft. Headquartered in Toronto, we have key offices in London, New York City, San Francisco, Montreal, Paris, Berlin, and Seoul.

About the Role

Our team is a fast-growing group of researchers and engineers focused on building reliable ML systems and pushing the boundaries of LLM inference efficiency. We develop techniques to improve how models execute in production, driving lower latency, higher throughput, and consistent quality across diverse workloads.

As an engineer on this team, you’ll work across the inference stack to improve core performance metrics by diving deep into model execution, identifying bottlenecks, and developing innovative optimizations. You’ll collaborate closely with modeling and systems teams to experiment, measure, and ship improvements that accelerate inference.

As the team evolves, you’ll have opportunities to build expertise in advanced performance techniques, including GPU/CUDA optimizations, kernel-level improvements, and model execution strategies for MoE and large-scale architectures.

We embrace a remote-friendly environment, strategically distributing teams based on interests, expertise, and time zones. The Model Efficiency team is concentrated in the EST and PST time zones—these are our preferred locations.

Requirements

  • 5+ years of experience writing high-performance, production-quality code
  • Strong programming skills in C++ or Python (Rust/Go also welcome)
  • Experience working with large language models and familiarity with the LLM inference ecosystem (e.g., vLLM, SGLang, etc.)
  • Ability to diagnose and resolve performance bottlenecks across the model execution stack
  • A strong bias for action—you ship fast, measure impact, and iterate

Nice to Have

  • GPU programming, CUDA, or low-level systems optimization
  • Language modeling with transformers (MoE, speculative decoding, KV-cache optimizations)
  • Scaling performance-critical distributed systems (e.g., computation, search, storage)

Benefits

  • A weekly lunch stipend of $75/£75 or equivalent in your local currency
  • Full health and dental benefits, including a separate budget for mental health
  • RRSP matching, 401K, or Pension Scheme
  • 100% Parental Leave top-up for up to 6 months for either parent
  • Annual enrichment benefits: Arts & culture, fitness/wellness, quality time, and a workspace improvement credit
  • Education & learning stipend for conferences, courses, and coaching
  • 6 weeks of paid vacation (30 working days)
  • Budget for traveling to other offices if remote, plus an annual company offsite

Work Environment

  • For those in the office: Daily lunch program, plenty of snacks, and regular community and social events.
  • For those not near an office: A co-working benefit to work alongside others in your city.
  • Everyone receives a $500 home office stipend to set up your workspace properly.

Similar jobs