Jobs · Nevada

AI Systems Engineer - DevOps& Observability - Senior

EY · Las Vegas, NV · 2 days ago
Hybrid$107k–$177k/yrFull-time

About the role

We are seeking an AI Systems Engineer to own the delivery, model-serving, routing, and observability layer of EY’s AI-native platform.

Responsibilities

  • Supports DevOps and delivery for AI workloads: build and operate the CI/CD/CV pipelines that ship AI services, agents, and runtime components, including automated build, test, continuous verification, release, and rollback, so AI workloads are delivered repeatably and safely into every environment.
  • Own governance and discovery for AI assets, including service catalog/registry (Artifactory/Nexus, Harbor), experiment tracking and model metadata (MLflow), upstream registries/mirrors (HuggingFace/NGC), CVE/SBOM scanning (Trivy), lineage contracts (OpenLineage), and license management.
  • Own resource and cost management, including quotas and rate limits, cost attribution and utilization (Apptio/OpenCost/Kubecost), so AI execution stays economically bounded and controllable per tenant and engagement.
  • Own the full observability stack, including metrics (Prometheus/Mimir), logs (Loki), traces (Tempo/Jaeger), dashboards (Grafana), LLM debugging and evaluation (LangSmith/Langfuse), and SLA/alert notifications.
  • Automate GitOps-based delivery and continuous verification; embedding quality, integrity, and cost gates into pipelines so releases are policy-compliant by default rather than by manual review.
  • Close the loop between delivery and observability by using telemetry, evaluation, and cost signals to drive deployment decisions, progressive rollout, and automated rollback of AI workloads.
  • Ensure cost and telemetry are identity-stamped and per-tenant, so consumption and behavior are attributable end-to-end, keeping FinOps and observability tied to the workloads that generate the load.

Requirements

  • Strong DevOps expertise: CI/CD/CV pipeline design, GitOps, continuous verification, and progressive/automated release and rollback for production workloads.
  • Deep expertise operating model-serving and inference systems (Ray, vLLM/Triton/NIM) on GPUs at production scale.
  • Strong observability skills: metrics, logs, traces, and OpenTelemetry.
  • Familiarity with model/artifact governance, registries, CVE scanning, and license/lineage tracking.
  • Comfortable operating across cloud, on-prem, edge, and air-gapped environments with consistent runtime and telemetry semantics.
  • Strong communicator able to explain runtime, cost, and observability tradeoffs to engineers, architects, and leadership.

Attributes for Success

  • Hands-on production ownership of AI or high-throughput services.
  • Hands-on expertise operating inference/model-serving frameworks (Ray Serve, vLLM, Triton, or NIM) on GPU infrastructure.
  • Strong experience with observability stacks (Prometheus, Grafana, Loki, Tempo/Jaeger) and OpenTelemetry.
  • Experience with API gateways and request routing (Envoy or equivalent), including streaming responses.
  • Experience with cost management / FinOps tooling (OpenCost, Kubecost, or equivalent) and quota/rate-limit enforcement.
  • Familiarity with model/artifact registries and supply-chain scanning (Harbor, MLflow, Trivy/SBOM).
  • Proven track record operating AI or service infrastructure under compliance, security, or regulatory constraints.
  • Ability to define clean ownership boundaries and consumption contracts with platform, trust, and data teams.

What We Offer

  • Cumulative compensation range: $106,900 to $176,500 in the US, $128,400 to $200,600 in the New York City Metro Area, Washington State and California (excluding Sacramento).
  • Flexible vacation policy.
  • Winter/Summer breaks.
  • Personal/Family Care.
  • Other leaves of absence.

Similar jobs