Lead Applied AI Site Reliability Engineer II - PxE A&A
Deloitte · Jersey City, NJ · Today
Hybrid$113k–$232k/yrFull-time
About the role
The Lead Applied AI Site Reliability Engineer II plays a critical role in ensuring the reliability, performance, and operational integrity of high-visibility products and platforms. This role requires a hands-on approach to managing cloud platform engineering, observability, and performance and reliability engineering, with a focus on operating AI and agentic workloads.
Responsibilities
- Outcome-Driven Accountability: Drive reliability, performance, and cost outcomes measured in service-level objectives and error budgets, not raw uptime.
- Technical Leadership and Advocacy: Advocate for production reliability and operability, set standards, and lead testing and operational tooling.
- Engineering Craftsmanship: Maintain operational integrity, build and operate production observability, and own SLOs and error budgets.
- Customer-Centric Engineering: Develop lean operational solutions through rapid experimentation to meet reliability needs of engineering teams and the business.
- Incremental and Iterative Delivery: Adopt a leaning-forward approach to reliability through incremental improvements and progressively refine SLOs.
- Cross-Functional Collaboration and Integration: Work collaboratively with cross-functional partners to set standards and integrate constraints.
- Advanced Technical Proficiency: Deep expertise in site reliability, cloud platform ownership, observability, performance and capacity engineering, chaos engineering, and cloud/AI cost engineering, with applied AI fluency.
- Domain Expertise: Quickly acquire domain knowledge of operated products and platforms, including their AI-infused production failure modes.
- Effective Communication and Influence: Communicate complex technical concepts clearly and inspire teammates and product teams through well-structured arguments and evidence.
- Engagement and Collaborative Co-Creation: Engage with product engineering teams at all levels, foster a culture of co-creation, and align diverse perspectives.
Requirements
- A bachelor's degree in computer science, software engineering, data science, machine learning, or related discipline.
- 6+ years of software engineering and site reliability engineering experience operating large-scale, distributed, cloud-native systems.
- Experience in defining and owning SLIs, SLOs, and SLAs; error budgets; incident command and on-call; building and operating production observability; environment integrity and drift prevention; and segregation-of-duties controls.
- 3+ years of experience with cloud-native engineering and cloud platform ownership on Azure, AWS, or GCP, including AI/ML services and container orchestration.
- Experience with reliability and operational standards, SLO discipline, runbooks, and performance and resilience budgets, including leading, mentoring, and guiding team members.
- Prior experience operating AI/ML and agentic workloads, their reliability failure modes, MLOps/LLMOps, and AI control plane.
- Experience with load and performance testing, chaos engineering, capacity planning, autoscaling, and cloud/AI cost engineering.
- Understanding of BCD, sequence/activity/state/entity relationship/data flow diagrams, OOP/OOD, data structures, algorithms, and AI-augmented spec-driven development.
- Experience with methodologies & tools such as XP, Lean, DevSecOps, SRE, ADO, GitHub, SonarQube, MLflow, and multi-agent orchestration tools.
Qualifications
- Excellent interpersonal and organizational skills, able to handle diverse situations, complex projects, and changing priorities.
- Ability to work in local office at least 3 days per week, with occasional travel.
Benefits
Not specified.
Pay
$113,100 to $232,300
Schedule
Not specified.