Jobs · Tennessee

AI Systems Engineer - AI Platforms - Manager

EY · Memphis, TN · 3 days ago
Hybrid$126k–$230k/yrFull-time

About the role

We are seeking AI Systems Engineers to build and operate the foundational substrate that powers EY’s AI-native platform. This role owns the infrastructure and cloud-native platform layers of the Hybrid AI Multi-Environment Runtime (HAI).

Responsibilities

  • Own the cluster & cloud-native platform: compute, Kubernetes and scheduling, cluster fabric/networking, multi-tenancy, and distributed compute.
  • Own the infrastructure foundation: Ubuntu/OS, BMC/bare-metal, DPU architecture, and NVAIE (GPU/Network/DCGM).
  • Stand up and manage EY Agentic AI environments across cloud (EKS/AKS/GKE), on-prem AI Factory (RKE2/NVAIE), edge (K3s), and air-gapped deployment modes.
  • Deliver foundational platform capabilities such as Infrastructure Management, Kubernetes & Scheduling, and Cluster Fabric Management.
  • Own cluster lifecycle, autoscaling, GPU pooling/virtualization, and multi-tenancy boundaries (vCluster/Crossplane/Karpenter).
  • Own secure execution and inference: Ray Serve, vLLM/NIM/Triton, and NVIDIA Dynamo.
  • Own cognitive and routing: Envoy AI Gateway, semantic routing (vLLM-SR), model/prompt selection, and streaming response handling.
  • Collaborate with DevOps Engineers on deployment and delivery of the platform itself.
  • Ensure the substrate is modular and swappable, so components can be replaced without rewriting consumers.

Requirements

Deep expertise in cloud-native platform engineering (Kubernetes, networking, multi-tenancy at scale) and low-level infrastructure and systems engineering (bare-metal, GPU, DPU).

Advanced understanding of how compute, networking, storage, and scheduling interact to form a reliable, portable substrate.

Comfortable operating across cloud, on-prem, edge, and air-gapped environments simultaneously, with a portability-first mindset.

Strong command of AI gateways, semantic routing, and model/prompt selection under latency and cost constraints.

A passion for ensuring reliability, repeatability, and reduction of operational toil through automation and infrastructure-as-code.

Strong communicator able to explain infrastructure tradeoffs to engineers, architects, and leadership.

Natural product-ownership orientation toward foundational platforms, measured by the downstream velocity and reliability they enable.

Skills and Attributes for Success

  • Expertise in cloud-native platform engineering (Kubernetes, networking, multi-tenancy at scale) and low-level infrastructure and systems engineering (bare-metal, GPU, DPU).
  • Advanced understanding of how compute, networking, storage, and scheduling interact to form a reliable, portable substrate.
  • Comfortable operating across cloud, on-prem, edge, and air-gapped environments simultaneously, with a portability-first mindset.
  • Strong command of AI gateways, semantic routing, and model/prompt selection under latency and cost constraints.
  • A passion for ensuring reliability, repeatability, and reduction of operational toil through automation and infrastructure-as-code.
  • Strong communicator able to explain infrastructure tradeoffs to engineers, architects, and leadership.
  • Natural product-ownership orientation toward foundational platforms, measured by the downstream velocity and reliability they enable.

Qualifications

  • Bachelor’s or Master’s degree in Computer Science or related technical field.
  • 8+ years building or operating enterprise infrastructure, cloud platforms, or large-scale Kubernetes environments, including hands-on systems depth.
  • Hands-on expertise with Kubernetes distributions (RKE2, EKS/AKS/GKE, K3s) and full cluster lifecycle management.
  • Deep experience with bare-metal, cloud, hybrid, on-prem, and ideally air-gapped deployment models.
  • Strong grounding in cluster networking (Cilium/service mesh/CNI), storage, and multi-tenancy isolation.
  • Experience with GPU infrastructure and scheduling (NVAIE/DCGM, GPU operators, virtualization/MIG).
  • Experience with infrastructure-as-code and GitOps tooling (Terraform/OpenTofu, Helm, ArgoCD, Crossplane).
  • Proven track record operating production infrastructure under compliance, security, or regulatory constraints.
  • Ability to work effectively with security, architecture, product, and delivery teams.

Benefits

The base salary range for this job in all geographic locations in the US is $125,500 to $230,200. The base salary range for New York City Metro Area, Washington State and California (excluding Sacramento) is $150,700 to $261,600. Individual salaries within those ranges are determined through a wide variety of factors including but not limited to education, experience, knowledge, skills and geography.

Pay

The base salary range for this job in all geographic locations in the US is $125,500 to $230,200. The base salary range for New York City Metro Area, Washington State and California (excluding Sacramento) is $150,700 to $261,600. Individual salaries within those ranges are determined through a wide variety of factors including but not limited to education, experience, knowledge, skills and geography.

Schedule

Our expectation is for most people in external, client serving roles to work together in person 40-60% of the time over the course of an engagement, project or year.

Similar jobs