Lead Applied AI Site Reliability Engineer II - PxE A&A
Deloitte · Hermitage, TN · Today
Hybrid$113k–$232k/yrFull-time
About the role
The Lead Applied AI Site Reliability Engineer II plays a critical role in ensuring the reliability, performance, and operational integrity of high-visibility products and platforms. This position requires a hands-on approach to managing production environments, setting production standards, and driving operational excellence.
Responsibilities
- Drive accountability for reliability, performance, and cost outcomes, measured in service-level objectives and error budgets.
- Operate products, platforms, and environments to meet their Service Level Indicators (SLIs) within budget, and track incident trends to prioritize work that improves reliability.
- Set production standards, lead observability, performance, and resilience testing, and operational tooling, and own the admission of systems into production.
- Maintain accountability for the operational integrity of production and pre-production environments, and for the production standards that systems are admitted against.
- Own Service Level Indicators (SLIs) and error budgets; build and operate production observability, including codified, version-controlled dashboards and SLO-driven, actionable alerting.
- Run performance, ambient-noise, and chaos testing to verify readiness; and guard environments against drift.
- Adopt a mindset that favors action and evidence over extensive planning, adopting a leaning-forward approach to reliability through incremental, measurable improvements.
- Work collaboratively with cross-functional partners to set production standards and integrate their constraints, co-defining service-level objectives and operational readiness.
- Establish reliability and operational standards, including SLO discipline, runbooks, and performance and resilience budgets, and guide team members in continuous improvement.
- Operate AI and agentic workloads reliably, including AI and Agentic SSDLC, delivering production operations with full automation from discovery to production to operations and all quality checks through the SSDLC lifecycle.
Qualifications
- Bachelor's degree in computer science, software engineering, data science, machine learning, or related discipline.
- 6+ years of software engineering and site reliability engineering experience operating large-scale, distributed, cloud-native systems in production.
- Experience in defining and owning Service Level Indicators (SLIs), Service Level Agreements (SLAs), error budgets, incident command, on-call, and building and operating production observability.
- Experience with cloud-native engineering and cloud platform ownership on any of the cloud hyperscalers such as Azure, AWS, or GCP.
- Experience with load and performance testing under simulated production traffic, chaos engineering, capacity planning, autoscaling, and cloud/AI cost engineering.
- Prior experience operating AI/ML and agentic workloads in production, including their reliability failure modes, MLOps/LLMOps, and the AI control plane.
- Experience with methodologies & tools such as XP, Lean, DevSecOps, SRE, ADO, GitHub, SonarQube, MLflow, and agentic AI frameworks.
- Ability to work in your local office at a minimum of 3 days per week.
Skills
- Python, Go, Bash, Java, C#/.NET, SQL/NoSQL, Kubernetes, Terraform, ArgoCD, CI/CD, observability stacks.
- Cloud platform ownership on Azure, AWS, or GCP, including AI/ML services like Azure OpenAI, AWS Bedrock, or Vertex AI.
- Container orchestration (Kubernetes, Docker), infrastructure-as-code, networking, and multi-environment management.
- Experience with reliability and operational standards, including SLO discipline, runbooks, and performance and resilience budgets.
- Experience with load and performance testing, chaos engineering, capacity planning, autoscaling, and cloud/AI cost engineering.
- Understanding of Business Context Diagrams (BCD), sequence/activity/state/entity relationship/data flow diagrams, OOP/OOD, data structures, algorithms, and code instrumentations, and AI-augmented spec-driven development.
- Experience with methodologies & tools such as XP, Lean, DevSecOps, SRE, ADO, GitHub, SonarQube, MLflow, and agentic AI frameworks.
Benefits
Deloitte offers a comprehensive benefits package including health insurance, retirement plans, and paid time off.
Pay
$113,100 to $232,300 annually.
Schedule
Flexible schedule with a minimum requirement of 3 days per week in the office.
Location
Commutable distance to select locations.
Equal Opportunity Employer
Deloitte is an equal opportunity employer committed to diversity and inclusion.