Lead Applied AI Site Reliability Engineer II - PxE A&A
Deloitte · New York, NY · Today
HybridEngineering$113k–$232k/yrFull-time
About the role
The Lead Applied AI Site Reliability Engineer II plays a critical role in ensuring the reliability, performance, and operational integrity of high-visibility products and platforms. This position requires a hands-on approach to managing production environments, setting production standards, and driving operational excellence.
Responsibilities
- Outcome-Driven Accountability: Drive reliability, performance, and cost outcomes measured in service-level objectives and error budgets, not raw uptime.
- Operational Excellence: Operate products, platforms, and environments to meet Service Level Objectives (SLOs) within budget, track incident trends, and prioritize work to improve reliability.
- Technical Leadership: Advocate for production reliability and operability, set production standards, and lead the design of observability, performance, and resilience testing.
- Engineering Craftsmanship: Maintain accountability for the operational integrity of production and pre-production environments, own SLOs and error budgets, and build and operate production observability.
- Customer-Centric Engineering: Develop lean operational solutions through rapid, inexpensive experimentation to meet the reliability needs of engineering teams and the business.
- Collaboration and Integration: Work collaboratively with cross-functional partners, set production standards, and integrate their constraints to ensure reliable, performant, and compliant operations.
- Advanced Technical Proficiency: Possess deep expertise in site reliability and modern production engineering, including cloud platform ownership, observability, performance and capacity engineering, chaos engineering, and cloud/AI cost engineering.
- Domain Expertise: Quickly acquire domain knowledge of products and platforms operated, translating reliability needs into service-level objectives, runbooks, and production tooling.
- Effective Communication and Influence: Exhibit exceptional communication skills, articulate complex technical concepts clearly, and inspire and influence teammates and product teams.
Qualifications
- A bachelor's degree in computer science, software engineering, data science, machine learning, or related discipline.
- 6+ years of software engineering and site reliability engineering experience operating large-scale, distributed, cloud-native systems.
- Experience in defining and owning Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Service Level Agreements (SLAs); error budgets; incident command and on-call; building and operating production observability; environment integrity and drift prevention; and segregation-of-duties controls.
- Experience with cloud-native engineering and cloud platform ownership on Azure, AWS, or GCP, including AI/ML services and container orchestration.
- Experience with reliability and operational standards, SLO discipline, runbooks, and performance and resilience budgets.
- Experience with AI/ML and agentic workloads, including their reliability failure modes, MLOps/LLMOps, and AI control plane.
- Experience with load and performance testing, chaos engineering, capacity planning, autoscaling, and cloud/AI cost engineering.
- Experience with methodologies and tools such as XP, Lean, DevSecOps, SRE, ADO, GitHub, SonarQube, MLflow, and multi-agent orchestration tools.
- Ability to work in the local office at least 3 days per week, with occasional travel.
Benefits
Comprehensive benefits package including health insurance, retirement plans, paid time off, and more.
Pay
$113,100 to $232,300 annually.
Schedule
Flexible schedule to accommodate the needs of the role and the business.