Jobs · Washington

Staff Product Manager - Observability

Lambda · Bellevue, WA · 2 wk ago
Hybrid$291k–$430k/yrFull-time

Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence.

This position requires presence in our Bellevue or San Francisco office location 4 days per week; Lambda's designated work from home day is currently Tuesday.

About The Role

A customer runs distributed training across 64 to 1,024+ GPUs (Graphics Processing Units). When a job slows down, customers need one question answered instantly: is it my code, NCCL (NVIDIA Collective Communications Library) configuration issue, or a bad InfiniBand link? Observability is how Lambda answers that question. It lets customers distinguish their code, their configuration, and our platform in minutes, and that transparency is a core reason teams trust Lambda with their largest training runs. This role makes that transparency real.

You will be the product manager who owns observability as a product across Lambda's cloud. That means everything customers see: GPU and cluster health, utilization, job-level telemetry, and InfiniBand fabric metrics. It also means the platform telemetry layer underneath, which powers Lambda's internal fleet operations and reliability work. You will partner daily with SRE (Site Reliability Engineering), fleet engineering, and the console and API (Application Programming Interface) teams to bring one coherent observability experience to customers and operators alike.

This role reports to the Head of Platform Product Management. Great product managers at Lambda are defined by three things: insight, influence, and execution. Insight means you look at the data, determine what it means for customers and business, and then figure out what to do about it. Influence means you take that idea and get others to want to buy into it; you win over engineers, designers, executives, and partners without relying on authority. Execution means you work with the right people to get the idea launched, then measure and iterate. We hire product managers who learn new domains fast and reason rigorously from evidence. Deep observability domain experience is a strong plus, but insight, influence, and execution are the bar.

Responsibilities

  • Own the Observability Roadmap: Define and drive the product strategy for observability across Lambda's cloud, from single On-Demand GPU Instances to 1-Click Clusters at 64 to 1,024+ GPU scale.
  • Turn Telemetry into Product: Decide which signals customers see, including GPU and cluster health, utilization, job-level telemetry, and InfiniBand fabric metrics, and how they see them in the console and the Lambda Cloud API.
  • Ship Diagnostics That Answer the First Question: Build experiences that let a customer quickly distinguish their own code or configuration issues from platform problems such as a failing node or a degraded InfiniBand link.
  • Work Backwards: Define and productize SLIs, SLOs, and SLAs with each product to help customers determine if their application or Lambda is causing issues, and if it's within expected tolerance.
  • Build the Platform Telemetry Layer: Partner with SRE and fleet engineering to define the internal telemetry that powers fleet operations and reliability, so one data foundation serves both customers and operators.
  • Live with Customers: Spend time with AI research labs, AI startups, and enterprise ML (machine learning) teams to understand how they debug distributed training and inference, and turn that signal into decisions about what to build next.
  • Win the Room: Bring engineers, designers, and executives along on your roadmap through clear writing and evidence, not authority.
  • Measure and Iterate: Define the success metrics for observability features, instrument adoption, and use the results to decide what ships next.
  • Deliver with the Console and API Teams: Drive observability capabilities through to launch across the console, the Lambda Cloud API, and cluster tooling, so customers can consume telemetry however they work.

Requirements

  • 7+ years of product management experience, including 3+ years on technical infrastructure, platform, or developer-facing products.
  • Ability to turn data and customer signal into a clear decision about what to build next, with examples where your insight changed a roadmap.
  • Experience winning over engineers, designers, executives, and customers without formal authority, and shipping products that required cross-team adoption.
  • Experience taking products from idea through launch, measurement, and iteration, with insights gained from post-launch metrics.
  • Comfort in deeply technical conversations about distributed systems, metrics pipelines, hardware health, and failure modes.
  • Strong writing skills to align teams with concise, clear documentation.
  • Ability to thrive in ambiguity, shape a new domain, and define iterative plans to move from current to desired state.

Nice to Have

  • Experience shipping monitoring or observability products, such as those in the Datadog, Grafana, or Prometheus class.
  • Worked with HPC (high performance computing) or distributed systems telemetry, including interconnect and fabric-level metrics.
  • Built developer tools or API-first products.
  • Hands-on exposure to distributed training stacks such as PyTorch with NCCL, or to GPU fleet tooling such as DCGM (Data Center GPU Manager).
  • Experience in a usage-based cloud infrastructure business.

Benefits

  • Generous cash & equity compensation.
  • Health, dental, and vision coverage for you and your dependents.
  • Wellness and commuter stipends for select roles.
  • 401k Plan with 2% company match (USA employees).
  • Flexible paid time off plan that we all actually use.

Pay

Compensation Range: $291K - $430K

Similar jobs