Jobs · Marketing · California

Principal Product Manager, AI Infrastructure and Orchestration

DataRobot · San Francisco, CA · 1 wk ago
MarketingFull-time

DataRobot delivers AI that maximizes impact and minimizes business risk. Our platform and applications integrate into core business processes so teams can develop, deliver, and govern AI at scale. We empower practitioners to deliver predictive and generative AI, and enable leaders to secure their AI assets. Organizations worldwide rely on DataRobot for AI that makes sense for their business — today and in the future.

We run agents and models in production for enterprises that cannot move their workloads to a public cloud: regulated industries, sovereign deployments, air-gapped and customer-managed clusters. The control plane is the layer that makes this possible. It deploys the agents, it deploys the models those agents call, and it decides how every workload is placed, scaled, isolated, routed to, and torn down across Kubernetes clusters and heterogeneous accelerators, on our cloud and on the customer's.

About the role

You will own the control plane layer as a product, focusing on the deployment and orchestration of agent and model workloads. This includes design reviews, API contracts, and production data analysis. Agents and models present a unique deployment challenge: agents are long-lived workloads with session state and unpredictable fan-out, calling models with uneven cost profiles, all competing for the same finite pool of accelerators. Your job is to allocate that pool correctly.

Responsibilities

  • Own the deployment and workload API, including resource model, lifecycle semantics, versioning, backward compatibility, and error behavior that customers integrate against.
  • Define placement and capacity strategies: how workloads land on nodes and accelerators, quota and priority across tenants, and behavior under contention.
  • Design scaling mechanisms: autoscaling signals, cold start and scale-to-zero economics, headroom policy, and cost-versus-latency trade-offs as customer-facing controls.
  • Manage the agent runtime: where agents run, their isolation, tool-call execution, and state persistence through restarts or evictions.
  • Oversee traffic and connectivity: ingress and routing for model and agent endpoints, request-aware load balancing, tenancy boundaries, and private connectivity into customer networks.
  • Implement governance and audit: track who deployed what, who invoked it, under which policy and access control, and ensure records survive customer audits.
  • Develop metering and packaging: measure inference, enforce quotas, attribute usage to tenants, and define pricing.
  • Ensure reliability: define SLOs, manage error budgets, and provide operational surfaces for platform engineers to diagnose issues without opening tickets.
  • Hold the roadmap for this layer across two engineering pods and align with product teams building on top of it.

Requirements

  • 6+ years in product management for infrastructure, developer platforms, or cloud services, with at least 3 years focused on Kubernetes-based or distributed systems products.
  • Principal candidates should have 9+ years of experience and a track record of owning a platform layer that other product teams built upon.
  • Deep technical understanding of GPU and accelerator behavior, including topology-aware placement, fractional and time-sliced sharing, MIG, device plugins, driver and container runtime plumbing, memory constraints, and utilization costs.
  • Deep technical understanding of Kubernetes: API server and scheduler, controllers and CRDs, operators, admission and RBAC, device plugins, resource requests and limits, node pools, and pod scheduling behavior.
  • Experience with multi-tenancy: isolation models, noisy neighbors, quota and fairness, and tenancy designs that pass customer security reviews.
  • Proven API product judgment: experience owning a public or platform API and managing the consequences of its contract.
  • Technical writing and prototyping as your default way to make a case, producing documents engineers will read, deep dives, public posts, API references, or working prototypes.
  • Comfort operating with matrixed engineering teams and no direct reports.
  • BS or MS in Computer Science or a closely related technical field, or equivalent hands-on experience as a software, platform, or infrastructure engineer.

Nice to have

  • Service networking depth: ingress and routing, load balancing under uneven request costs, DNS, TLS termination, private link connectivity, and network policy as a tenancy boundary.
  • Experience with modern serving stacks and their failure modes, such as vLLM, KV cache behavior, batching, and quantization trade-offs.
  • Knowledge of long-running and agentic workload patterns: session affinity, statefulness, tool-call fan-out, and sandboxed execution.
  • Experience with customer-managed, air-gapped, or sovereign deployments and their compliance constraints.
  • Regulated-industry experience where access control, secrets, and invocation-level auditability are product requirements.
  • CNCF or open-source contributions.

How we work with AI

We use AI tooling in the loop daily and expect the same from candidates. Strong candidates "build to think": creating prototypes to test APIs before specification, evaluation harnesses, agents wired against real services, or throwaway tools to answer roadmap questions quickly. Bring one or two things you built yourself—production-grade code is not required, but the ability to iterate without waiting on engineers is.

Scope of the role

This role owns workload and agent runtime orchestration: placement, lifecycle, scaling, and isolation of agent and model workloads on Kubernetes and accelerators. It does not cover agent coordination and authoring (planner/router logic, memory strategy, multi-agent handoff, builder surfaces), workflow orchestration (Airflow/Temporal), model quality, applied research, or cluster fabric work (InfiniBand/RoCE tuning). There are no direct reports; scope is defined by technical surface area and influence across engineering pods.

Benefits

  • Medical, Dental & Vision Insurance
  • Flexible Time Off Program
  • Paid Holidays
  • Paid Parental Leave
  • Global Employee Assistance Program (EAP)
  • Additional location-specific benefits as required by local legal requirements

DataRobot Operating Principles

  • Wow Our Customers
  • Set High Standards
  • Be Better Than Yesterday
  • Be Rigorous
  • Assume Positive Intent
  • Have the Tough Conversations
  • Be Better Together
  • Debate, Decide, Commit
  • Deliver Results
  • Overcommunicate

Similar jobs