AI Systems Engineer - AI Platforms - Manager
About the role
We are seeking AI Systems Engineers to build and operate the foundational substrate that powers EY’s AI-native platform. This role owns the infrastructure and cloud-native platform layers of the Hybrid AI Multi-Environment Runtime (HAI), from bare-metal and GPU infrastructure through Kubernetes, cluster fabric, and multi-tenant scaling.
Responsibilities
- Own the cluster & cloud-native platform: compute, Kubernetes and scheduling, cluster fabric/networking, multi-tenancy, and distributed compute, as the substrate for Agentic AI workflows and tooling.
- Own the infrastructure foundation: Ubuntu/OS, BMC/bare-metal, DPU architecture, and NVAIE (GPU/Network/DCGM), ensuring the physical and virtual bedrock is provisioned, patched, and production-ready.
- Stand up and manage EY Agentic AI environments across cloud (EKS/AKS/GKE), on-prem AI Factory (RKE2/NVAIE), edge (K3s), and air-gapped deployment modes, maintaining one consistent stack contract across all targets.
- Deliver foundational platform capabilities such as Infrastructure Management, Kubernetes & Scheduling, and Cluster Fabric Management, so downstream runtime, data, and execution services can run safely and consistently.
- Own cluster lifecycle, autoscaling, GPU pooling/virtualization, and multi-tenancy boundaries (vCluster/Crossplane/Karpenter), providing isolated, elastic capacity per tenant and engagement.
- Own secure execution and inference: Ray Serve, vLLM/NIM/Triton, and NVIDIA Dynamo, with sandboxed execution (gVisor/Firecracker for hosted, NVIDIA OpenShell/vNode for on-prem) for isolated, safe model execution.
- Own cognitive and routing: Envoy AI Gateway, semantic routing (vLLM-SR), model/prompt selection, and streaming response handling — directing each request to the right model under the right constraints.
- Collaborate with DevOps Engineers on deployment and delivery of the platform itself: CI/CD/CV (ArgoCD), infrastructure-as-code / GitOps (Helm/OpenTofu), so environments are reproducible and drift-free.
- Own backup, disaster recovery, and cross-environment replication for high availability (Velero, CloudNativePG, Cilium ClusterMesh), along with patching and platform supply-chain hygiene.
- Ensure the substrate is modular and swappable, so components can be replaced without rewriting consumers, minimizing vendor lock-in while preserving the stack contract.
Requirements
Deep expertise in cloud-native platform engineering (Kubernetes, networking, multi-tenancy at scale) and low-level infrastructure and systems engineering (bare-metal, GPU, DPU).
Advanced understanding of how compute, networking, storage, and scheduling interact to form a reliable, portable substrate.
Comfortable operating across cloud, on-prem, edge, and air-gapped environments simultaneously, with a portability-first mindset.
Strong command of AI gateways, semantic routing, and model/prompt selection under latency and cost constraints.
A passion for ensuring reliability, repeatability, and reduction of operational toil through automation and infrastructure-as-code.
Strong communicator able to explain infrastructure tradeoffs to engineers, architects, and leadership.
Natural product-ownership orientation toward foundational platforms, measured by the downstream velocity and reliability they enable.
Qualifications
- Bachelor’s or Master’s degree in Computer Science or related technical field.
- 8+ years building or operating enterprise infrastructure, cloud platforms, or large-scale Kubernetes environments, including hands-on systems depth.
- Hands-on expertise with Kubernetes distributions (RKE2, EKS/AKS/GKE, K3s) and full cluster lifecycle management.
- Deep experience with bare-metal, cloud, hybrid, on-prem, and ideally air-gapped deployment models.
- Strong grounding in cluster networking (Cilium/service mesh/CNI), storage, and multi-tenancy isolation.
- Experience with GPU infrastructure and scheduling (NVAIE/DCGM, GPU operators, virtualization/MIG).
- Experience with infrastructure-as-code and GitOps tooling (Terraform/OpenTofu, Helm, ArgoCD, Crossplane).
- Proven track record operating production infrastructure under compliance, security, or regulatory constraints.
- Ability to work effectively with security, architecture, product, and delivery teams.
- Ideally, you’ll also have Experience supporting AI/ML and GPU-intensive workloads (GPU virtualization, MIG, AI workload scheduling).
- Experience with multi-tenant platforms and tenancy isolation at scale (e.g., vCluster).
- Experience with disaster recovery, backup, and cross-environment replication tiered by RPO/RTO.
- Experience with distributed compute and batch orchestration (Kueue, Slurm/Slinky, Ray, OpenMPI).
- Background in platform or SRE roles where success is measured by downstream team velocity, reliability, and portability.
- Exposure to regulated delivery environments (financial services, tax, healthcare, risk).
Skills and Attributes
- Deep expertise in cloud-native platform engineering (Kubernetes, networking, multi-tenancy at scale) and low-level infrastructure and systems engineering (bare-metal, GPU, DPU).
- Advanced understanding of how compute, networking, storage, and scheduling interact to form a reliable, portable substrate.
- Comfortable operating across cloud, on-prem, edge, and air-gapped environments simultaneously, with a portability-first mindset.
- Strong command of AI gateways, semantic routing, and model/prompt selection under latency and cost constraints.
- A passion for ensuring reliability, repeatability, and reduction of operational toil through automation and infrastructure-as-code.
- Strong communicator able to explain infrastructure tradeoffs to engineers, architects, and leadership.
- Natural product-ownership orientation toward foundational platforms, measured by the downstream velocity and reliability they enable.
Pay
$125,500 to $230,200
Schedule
Hybrid model with in-person work expected 40-60% of the time over the course of an engagement, project or year.
Benefits
Comprehensive compensation and benefits package including medical and dental coverage, pension and 401(k) plans, and a wide range of paid time off options.