Jobs · New Jersey

Lead Applied AI Site Reliability Engineer II - PxE ERM

Deloitte · Jersey City, NJ · Today
Hybrid$113k–$232k/yrFull-time

About the role

The Lead Applied AI Site Reliability Engineer II plays a pivotal role in ensuring the reliability, performance, and operational integrity of high-visibility products and platforms. This role requires a hands-on approach to maintaining production safety, efficiency, and cost-effectiveness.

Responsibilities

  • Drive accountability for reliability, performance, and cost outcomes, measured in service-level objectives and error budgets.
  • Operate products, platforms, and environments to meet Service Level Indicators (SLIs) within budget, and track incident trends and toil to prioritize work that improves reliability.
  • Set production standards, lead observability, performance, and resilience testing, and operational tooling, and own the admission of systems into production.
  • Maintain accountability for the operational integrity of production and pre-production environments, and for the production standards that systems are admitted against.
  • Adopt a mindset that favors action and evidence over extensive planning, adopting a leaning-forward approach to reliability through incremental, measurable improvements.
  • Work collaboratively with cross-functional partners to set production standards and integrate their constraints, ensuring the reliable, performant, and compliant path is the operative path.
  • Establish reliability and operational standards, including SLO discipline, runbooks, and performance and resilience budgets, and guide team members in continuous improvement.
  • Operate AI and agentic workloads reliably, including AI and Agentic SSDLC, and deliver production operations with full automation from discovery to production to operations and all quality checks through the SSDLC lifecycle.

Requirements

  • A bachelor's degree in computer science, software engineering, data science, machine learning, or related discipline.
  • 6+ years of software engineering and site reliability engineering experience operating large-scale, distributed, cloud-native systems in production.
  • Experience in defining and owning SLIs, SLOs, and SLAs; error budgets; incident command and on-call; building and operating production observability; environment integrity and drift prevention; and segregation-of-duties controls.
  • Experience with cloud-native engineering and cloud platform ownership on any of the cloud hyperscalers such as Azure, AWS, or GCP, including AI/ML services and container orchestration.
  • Experience with reliability and operational standards, including SLO discipline, runbooks, and performance and resilience budgets.
  • Prior experience operating AI/ML and agentic workloads in production, including their reliability failure modes, MLOps/LLMOps, and the AI control plane.
  • Experience with load and performance testing, chaos engineering, capacity planning, autoscaling, and cloud/AI cost engineering.
  • Experience with methodologies & tools such as XP, Lean, DevSecOps, SRE, ADO, GitHub, SonarQube, MLflow, and agentic AI frameworks.

Qualifications

  • Excellent interpersonal and organizational skills.
  • Ability to handle diverse situations, complex projects, and changing priorities.
  • Passion, empathy, and care in handling diverse situations.

Skills

  • Python, Go, Bash, Java, C#/.NET, SQL/NoSQL, Kubernetes, Terraform, ArgoCD, CI/CD, observability stacks.
  • Cloud-native engineering and cloud platform ownership on Azure, AWS, or GCP.
  • Establishing reliability and operational standards.
  • Operating AI/ML and agentic workloads in production.
  • Load and performance testing, chaos engineering, capacity planning, autoscaling, and cloud/AI cost engineering.
  • XP, Lean, DevSecOps, SRE, ADO, GitHub, SonarQube, MLflow, and agentic AI frameworks.

Benefits

Not specified.

Pay

$113,100 to $232,300 annually.

Schedule

Not specified.

Similar jobs