Principal Product Manager, AI Infrastructure and Orchestration
DataRobot delivers AI that maximizes impact and minimizes business risk. Our platform and applications integrate into core business processes so teams can develop, deliver, and govern AI at scale. We empower practitioners to deliver predictive and generative AI, and enable leaders to secure their AI assets. Organizations worldwide rely on DataRobot for AI that makes sense for their business — today and in the future.
We run agents and models in production for enterprises that cannot move their workloads to a public cloud: regulated industries, sovereign deployments, air-gapped and customer-managed clusters. The control plane is the layer that makes this possible. It deploys the agents, it deploys the models those agents call, and it decides how every workload is placed, scaled, isolated, routed to, and torn down across Kubernetes clusters and heterogeneous accelerators, on our cloud and on the customer's.
About the role
You will own the control plane layer as a product, focusing on the deployment and orchestration of agent and model workloads. This involves design reviews, API contracts, and production data analysis to ensure efficient placement, scaling, and isolation of workloads on Kubernetes and accelerators.
Responsibilities
- Own the deployment and workload API, including resource model, lifecycle semantics, versioning, backward compatibility, and error behavior.
- Define placement and capacity strategies: how workloads land on nodes and accelerators, quota and priority across tenants, and behavior under contention.
- Manage scaling: autoscaling signals, cold start and scale-to-zero economics, headroom policy, and cost-versus-latency trade-offs as customer-facing controls.
- Oversee the agent runtime: where and how long agents run, isolation, tool call execution, and state survival during restarts or evictions.
- Handle traffic and connectivity: ingress and routing for model and agent endpoints, request-aware load balancing, tenancy boundaries, and private connectivity into customer networks.
- Ensure governance and audit: track deployments, invocations, policies, access control, and audit records.
- Define metering and packaging: how inference is measured, quota'd, attributed to tenants, and priced.
- Maintain reliability: establish SLOs, error budgets, and operational surfaces for diagnosing degraded deployments.
- Hold the roadmap for this layer across two engineering pods and align with product teams building on top of it.
Requirements
- 6+ years in product management for infrastructure, developer platforms, or cloud services, with at least 3 years on Kubernetes-based or distributed systems products.
- Principal candidates should have 9+ years and experience with a platform layer other product teams built on.
- Deep technical understanding of GPU and accelerator behavior: topology-aware placement, fractional/time-sliced sharing, MIG, device plugins, memory constraints, and utilization costs.
- Deep technical understanding of Kubernetes: API server, scheduler, controllers, CRDs, operators, admission/RBAC, device plugins, resource requests/limits, node pools, and pod scheduling.
- Multi-tenancy experience: isolation models, noisy neighbors, quota, fairness, and tenancy designs for security reviews.
- API product judgment: experience owning a public or platform API and managing its contract consequences.
- Technical writing and prototyping as your default way to make a case (e.g., docs, deep dives, API references, or working prototypes).
- Comfort operating with matrixed engineering teams and no direct reports.
- BS/MS in Computer Science or related field, or equivalent hands-on experience as a software/platform/infrastructure engineer.
Nice to have
- Service networking depth: ingress/routing, load balancing under uneven request costs, DNS, TLS termination, private link connectivity, and network policy as a tenancy boundary.
- Modern serving stacks and failure modes: vLLM or similar, KV cache behavior, batching, quantization trade-offs.
- Long-running and agentic workload patterns: session affinity, statefulness, tool-call fan-out, sandboxed execution.
- Customer-managed, air-gapped, or sovereign deployments and compliance constraints.
- Regulated-industry experience with access control, secrets, and invocation-level auditability.
- CNCF or open-source contributions.
How we work with AI
We use AI tooling daily and expect candidates to do the same. Strong candidates build to think: prototypes for API exploration, evaluation harnesses, agents wired to real services, or tools to answer roadmap questions quickly. Bring examples of self-built projects—production-grade code isn’t required, but the ability to iterate without engineering support is.
Scope of the role
This role owns workload and agent runtime orchestration: placement, lifecycle, scaling, and isolation of agent/model workloads on Kubernetes and accelerators. It does not cover:
- Agent coordination/authoring (planner/router logic, memory strategy, multi-agent handoff, builder surfaces).
- Workflow orchestration (Airflow/Temporal).
- Model quality/applied research.
- Cluster fabric work (InfiniBand/RoCE tuning for distributed training).
There are no direct reports; scope is defined by technical surface area and influence across engineering pods.
Benefits
- Medical, dental, and vision insurance.
- Flexible time off program and paid holidays.
- Paid parental leave.
- Global Employee Assistance Program (EAP).