AI Systems Engineer - DevOps& Observability - Senior
EY · Irvine, CA · 2 days ago
HybridInformation Technology$107k–$177k/yrFull-time
About the role
We are seeking an AI Systems Engineer to own the delivery, model-serving, routing, and observability layer of EY’s AI-native platform. This role advises how AI services and agents are built and deployed, how models execute, how requests are routed to them, how AI assets are catalogued and governed, how consumption is measured and bounded, and how the entire platform is observed across cloud, on-prem, edge, and air-gapped environments.
Responsibilities
- Supports DevOps and delivery for AI workloads: build and operate the CI/CD/CV pipelines that ship AI services, agents, and runtime components, including automated build, test, continuous verification, release, and rollback, so AI workloads are delivered repeatably and safely into every environment.
- Owns governance and discovery for AI assets, including service catalog/registry (Artifactory/Nexus, Harbor), experiment tracking and model metadata (MLflow), upstream registries/mirrors (HuggingFace/NGC), CVE/SBOM scanning (Trivy), lineage contracts (OpenLineage), and license management.
- Owns resource and cost management, including quotas and rate limits, cost attribution and utilization (Apptio/OpenCost/Kubecost), so AI execution stays economically bounded and controllable per tenant and engagement.
- Owns the full observability stack, including metrics (Prometheus/Mimir), logs (Loki), traces (Tempo/Jaeger), dashboards (Grafana), LLM debugging and evaluation (LangSmith/Langfuse), and SLA/alert notifications.
- Automates GitOps-based delivery and continuous verification; embedding quality, integrity, and cost gates into pipelines so releases are policy-compliant by default rather than by manual review.
- Closes the loop between delivery and observability by using telemetry, evaluation, and cost signals to drive deployment decisions, progressive rollout, and automated rollback of AI workloads.
- Ensures cost and telemetry are identity-stamped and per-tenant, so consumption and behavior are attributable end-to-end, keeping FinOps and observability tied to the workloads that generate the load.
Requirements
- Strong DevOps expertise: CI/CD/CV pipeline design, GitOps, continuous verification, and progressive/automated release and rollback for production workloads.
- Deep expertise operating model-serving and inference systems (Ray, vLLM/Triton/NIM) on GPUs at production scale.
- Strong hands-on DevOps experience, including CI/CD/CV pipelines and GitOps tooling (ArgoCD, Helm, GitHub Actions/GitLab CI, or equivalents) for automated build, test, release, and rollback.
- Hands-on expertise operating inference/model-serving frameworks (Ray Serve, vLLM, Triton, or NIM) on GPU infrastructure.
- Strong experience with observability stacks (Prometheus, Grafana, Loki, Tempo/Jaeger) and OpenTelemetry.
- Experience with API gateways and request routing (Envoy or equivalent), including streaming responses.
- Experience with cost management / FinOps tooling (OpenCost, Kubecost, or equivalent) and quota/rate-limit enforcement.
- Familiarity with model/artifact registries and supply-chain scanning (Harbor, MLflow, Trivy/SBOM).
- Proven track record operating AI or service infrastructure under compliance, security, or regulatory constraints.
Attributes for Success
- Comfortable operating across cloud, on-prem, edge, and air-gapped environments with consistent runtime and telemetry semantics.
- Strong communicator able to explain runtime, cost, and observability tradeoffs to engineers, architects, and leadership.