Jobs · Engineering · Texas

Lead Applied AI Site Reliability Engineer II - PxE A&A

Deloitte · Dallas, TX · Today
HybridEngineering$113k–$232k/yrFull-time

About the role

The Lead Applied AI Site Reliability Engineer II plays a crucial role in ensuring the reliability, performance, and operational integrity of high-visibility products and platforms. This position requires a hands-on approach to maintaining production safety, efficiency, and cost-effectiveness.

Responsibilities

  • Drive accountability for reliability, performance, and cost outcomes, measured in service-level objectives and error budgets, not raw uptime.
  • Operate products, platforms, and environments to meet Service Level Indicators (SLIs) within budget, track incident trends, and prioritize work to improve reliability.
  • Set production standards, lead observability, performance, and resilience testing, and operational tooling, and own the admission of systems into production.
  • Maintain accountability for the operational integrity of production and pre-production environments, and for the production standards systems are admitted against.
  • Own Service Level Indicators (SLOs) and error budgets; build and operate production observability, including codified, version-controlled dashboards and SLO-driven, actionable alerting; run performance, ambient-noise, and chaos testing; and guard environments against drift.
  • Collaborate with diverse teams to develop lean operational solutions through rapid, inexpensive experimentation to meet reliability needs.
  • Work collaboratively with cross-functional partners to set production standards and integrate their constraints, co-defining service-level objectives and operational readiness.
  • Establish reliability and operational standards, including SLO discipline, runbooks, and performance and resilience budgets, and guide team members in continuous improvement.
  • Operate AI and agentic workloads reliably, including AI and Agentic SSDLC, delivering production operations with full automation from discovery to production to operations and all quality checks through the SSDLC lifecycle.

Qualifications

  • Bachelor's degree in computer science, software engineering, data science, machine learning, or related discipline.
  • 6+ years of software engineering and site reliability engineering experience operating large-scale, distributed, cloud-native systems in production.
  • Experience in defining and owning Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Service Level Agreements (SLAs); error budgets; incident command and on-call; building and operating production observability; environment integrity and drift prevention; and segregation-of-duties controls.
  • Experience with cloud-native engineering and cloud platform ownership on any of the cloud hyperscalers such as Azure, AWS, or GCP.
  • Experience with load and performance testing under simulated production traffic, chaos engineering, capacity planning, autoscaling, and cloud/AI cost engineering.
  • Prior experience operating AI/ML and agentic workloads in production, including their reliability failure modes, MLOps/LLMOps, and the AI control plane.
  • Experience with methodologies & tools such as XP, Lean, DevSecOps, SRE, ADO, GitHub, SonarQube, MLflow, and agentic AI frameworks.
  • Ability to work in your local office at a minimum of 3 days per week.

Skills

  • Python, Go, Bash, Java, C#/.NET, SQL/NoSQL, Kubernetes, Terraform, ArgoCD, CI/CD, observability stacks.
  • Cloud platform ownership on Azure, AWS, or GCP, including AI/ML services and container orchestration.
  • Experience with load and performance testing, chaos engineering, capacity planning, autoscaling, and cloud/AI cost engineering.
  • Understanding of Business Context Diagrams (BCD), sequence/activity/state/entity relationship/data flow diagrams, OOP/OOD, data structures, algorithms, and code instrumentations, and AI-augmented spec-driven development.
  • Experience with XP, Lean, DevSecOps, SRE, ADO, GitHub, SonarQube, MLflow, and agentic AI frameworks.

Benefits

Not specified.

Pay

$113,100 to $232,300

Schedule

Not specified.

Similar jobs