Applied AI SRE III - PxE GPS
Deloitte · Morristown, NJ · 1 mo ago
Hybrid$103k–$211k/yrFull-time
About the role
The Applied AI Site Reliability Engineer III plays a critical role in ensuring the reliability, performance, and operational integrity of high-visibility products and platforms. This position requires a hands-on approach to maintaining production safety, efficiency, and cost-effectiveness.
Responsibilities
- Drive accountability for reliability, performance, and cost outcomes, measured in service-level objectives and error budgets.
- Operate products, platforms, and environments to meet Service Level Indicators (SLIs) within budget, track incident trends, and prioritize work to improve reliability.
- Lead the design of observability, performance, and resilience testing, and operational tooling, and own the admission of systems into production.
- Maintain accountability for the operational integrity of production and pre-production environments, and for the production standards that systems are admitted against.
- Adopt a leaning-forward approach to reliability, progressively refining SLOs and implementing resilience testing, rather than big-bang interventions.
- Work collaboratively with cross-functional partners to uphold production standards and integrate their constraints.
- Develop lean operational solutions through rapid, inexpensive experimentation to meet reliability needs of engineering teams and the business.
- Translate reliability needs, reference architectures, and operational requirements into service-level objectives, runbooks, and production tooling.
- Operate high-quality, resilient platforms and products at scale, leveraging advanced technical proficiency in site reliability and modern production engineering.
Requirements
- A bachelor's degree in computer science, software engineering, data science, machine learning, or related discipline.
- 5+ years of software engineering and site reliability engineering experience operating large-scale, distributed, cloud-native systems in production.
- Experience defining and owning SLIs, SLOs, and SLAs; error budgets; incident command and on-call; building and operating production observability; environment integrity and drift prevention; and segregation-of-duties controls.
- 3+ years of experience with cloud-native engineering and cloud platform ownership on any of the cloud hyperscalers such as Azure, AWS, or GCP.
- Prior experience operating AI/ML and agentic workloads in production, including their reliability failure modes, MLOps/LLMOps, and the AI control plane.
- Experience with load and performance testing, chaos engineering, capacity planning, autoscaling, and cloud/AI cost engineering.
- Understanding of Business Context Diagrams (BCD), sequence/activity/state/entity relationship/data flow diagrams, OOP/OOD, data structures, algorithms, and code instrumentations, and AI-augmented spec-driven development.
- Experience using methodologies & tools such as XP, Lean, DevSecOps, SRE, ADO, GitHub, SonarQube, MLflow, and agentic AI frameworks.
Qualifications
- Legal authorization to work in the United States without the need for employer sponsorship.
- U.S. Citizenship or Green Card holder.
Skills
- Excellent interpersonal and organizational skills.
- Advanced technical proficiency in site reliability and modern production engineering.
- Expertise in cloud platform ownership, observability, performance and capacity engineering, chaos engineering, and cloud/AI cost engineering.
- Applied AI fluency to operate AI and agentic workloads reliably.
Benefits
Details not specified.
Pay
$102,500 to $210,600 annually.
Schedule
Minimum 3 days per week in the local office, with occasional travel up to 10%.