Jobs · Tennessee

Applied AI SRE III - PxE GPS

Deloitte · Hermitage, TN · 1 mo ago
Hybrid$103k–$211k/yrFull-time

About the role

The Applied AI Site Reliability Engineer III plays a crucial role in ensuring the reliability, performance, and operational integrity of high-visibility products and platforms. This position requires a hands-on approach to managing cloud platform engineering, observability, and performance and reliability engineering, with a focus on operating AI and agentic workloads.

Responsibilities

  • Embrace and drive a culture of accountability for reliability, performance, and cost outcomes, measured in service-level objectives and error budgets.
  • Operate the products, platforms, and environments you support to meet their Service Level Objectives (SLOs) within budget, and track incident trends and toil to prioritize work that most improves reliability.
  • Own SLOs and error budgets; build and operate production observability-codified, version-controlled dashboards and SLO-driven, actionable alerting that detects before impact, plus the feedback loop into engineering.
  • Run performance, ambient-noise, and chaos testing to verify readiness; guard environments against drift.
  • Maintain accountability for the operational integrity of production and pre-production environments, and for the production standards that systems are admitted against.
  • Adopt a mindset that favors action and evidence over extensive planning, adopting a leaning-forward approach to navigate complexity and uncertainty, hardening reliability through incremental, measurable improvements.
  • Work collaboratively with empowered, cross-functional partners: engineering, platform engineering, security and risk, data governance, and engineering leadership and architecture.
  • Develop lean operational solutions through rapid, inexpensive experimentation to meet the reliability needs of the engineering teams and the business.
  • Translate reliability needs, reference architectures, and operational requirements into service-level objectives, runbooks, and production tooling.

Requirements

  • A bachelor's degree in computer science, software engineering, data science, machine learning, or related discipline.
  • Experience operating large-scale, distributed, cloud-native systems in production, with experience in Python, Go, Bash, Java, C#, SQL/NoSQL, Kubernetes, Terraform, ArgoCD, CI/CD, and observability stacks.
  • Experience defining and owning SLIs, SLOs, and SLAs; error budgets; incident command and on-call; building and operating production observability; environment integrity and drift prevention; and segregation-of-duties controls.
  • Experience with cloud-native engineering and cloud platform ownership on Azure, AWS, or GCP, including AI/ML services and container orchestration.
  • Prior experience operating AI/ML and agentic workloads in production, including their reliability failure modes, MLOps/LLMOps, and the AI control plane.
  • Experience with load and performance testing, chaos engineering, capacity planning, autoscaling, and cloud/AI cost engineering.
  • Understanding of Business Context Diagrams, sequence/activity/state/entity relationship/data flow diagrams, OOP/OOD, data structures, algorithms, and code instrumentations, and AI-augmented spec-driven development.
  • Experience using methodologies & tools such as XP, Lean, DevSecOps, SRE, ADO, GitHub, SonarQube, MLflow, and multi-agent orchestration tools.

Qualifications

  • Legal authorization to work in the United States without the need for employer sponsorship.
  • U.S. Citizenship or Green Card holder.

Skills

  • Excellent interpersonal and organizational skills.
  • Strong track record in operating high-quality, resilient systems at scale.
  • Expertise in site reliability and modern production engineering, cloud platform ownership, observability, performance and capacity engineering, chaos engineering, and cloud/AI cost engineering.
  • Fluency in applied AI to operate AI and agentic workloads reliably.

Benefits

N/A

Pay

$102,500 to $210,600 annually

Schedule

Minimum of 3 days per week in the office, with occasional travel required.

Similar jobs