Lead Applied AI Site Reliability Engineer II - PxE A&A
Deloitte · Tampa, FL · Today
Hybrid$113k–$232k/yrFull-time
About the role
The Lead Applied AI Site Reliability Engineer II role is pivotal in ensuring the reliability, performance, and operational integrity of high-visibility products and platforms. This role requires a hands-on approach to maintaining production safety, performance, and cost-effectiveness.
Responsibilities
- Outcome-Driven Accountability: Drive reliability, performance, and cost outcomes measured in service-level objectives and error budgets, not raw uptime.
- Technical Leadership and Advocacy: Advocate for production reliability and operability, set standards, and own the admission of systems into production.
- Engineering Craftsmanship: Maintain accountability for production and pre-production environments, own SLOs and error budgets, and build and operate production observability.
- Customer-Centric Engineering: Develop lean operational solutions through rapid, inexpensive experimentation to meet reliability needs of engineering teams and the business.
- Cross-Functional Collaboration and Integration: Work collaboratively with cross-functional partners to set production standards and integrate constraints.
- Advanced Technical Proficiency: Possess deep expertise in site reliability and modern production engineering, including cloud platform ownership, observability, performance and capacity engineering, chaos engineering, and cloud/AI cost engineering.
- Domain Expertise: Quickly acquire domain knowledge of operated products and platforms, including their distinct production failure modes.
- Effective Communication and Influence: Exhibit exceptional communication skills, articulate complex technical concepts clearly, and inspire and influence teammates and product teams.
- Engagement and Collaborative Co-Creation: Engage and collaborate with product engineering teams, build and maintain constructive relationships, and foster a culture of co-creation.
Qualifications
- A bachelor's degree in computer science, software engineering, data science, machine learning, or related discipline.
- 6+ years of software engineering and site reliability engineering experience operating large-scale, distributed, cloud-native systems in production.
- Experience in defining and owning SLIs, SLOs, and SLAs; error budgets; incident command and on-call; building and operating production observability; environment integrity and drift prevention; and segregation-of-duties controls.
- Experience with cloud-native engineering and cloud platform ownership on any of the cloud hyperscalers such as Azure, AWS, or GCP.
- Experience with load and performance testing under simulated production traffic, chaos engineering, capacity planning, autoscaling, and cloud/AI cost engineering.
- Prior experience operating AI/ML and agentic workloads in production, including their reliability failure modes, MLOps/LLMOps, and AI control plane.
- Experience with methodologies & tools such as XP, Lean, DevSecOps, SRE, ADO, GitHub, SonarQube, MLflow, and agentic AI frameworks.
- Ability to work in your local office at a minimum of 3 days per week.