Jobs · New Jersey

Applied AI SRE III - PxE GPS

Deloitte · Jersey City, NJ · 1 mo ago
Hybrid$103k–$211k/yrFull-time

About the role

The Applied AI Site Reliability Engineer III plays a crucial role in ensuring the reliability, performance, and operational integrity of high-visibility products and platforms. This position requires a hands-on approach to managing cloud platform engineering, observability, and performance and reliability engineering, with a focus on operating AI and agentic workloads.

Responsibilities

  • Drive accountability for reliability, performance, and cost outcomes, measured in service-level objectives and error budgets.
  • Operate products, platforms, and environments to meet their Service Level Objectives (SLOs) within budget, and track incident trends and toil to prioritize work that most improves reliability.
  • Ensure systems are admissible, performant, safe to run, and can degrade gracefully when failure occurs.
  • Lead the design of observability, performance and resilience testing, and operational tooling, and own the admission of systems into production.
  • Maintain accountability for the operational integrity of production and pre-production environments, and for the production standards that systems are admitted against.
  • Own Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Service Level Agreements (SLAs); manage error budgets; build and operate production observability; and run performance, ambient-noise, and chaos testing to verify readiness.
  • Admit systems into production with confidence, adhering to segregation-of-duties controls, and partnering with security and risk on control objectives.
  • Develop lean operational solutions through rapid, inexpensive experimentation to meet the reliability needs of the engineering teams and the business.
  • Work collaboratively with cross-functional partners to uphold production standards and integrate their constraints.
  • Adopt a mindset that favors action and evidence over extensive planning, adopting a leaning-forward approach to reliability through incremental, measurable improvements.

Qualifications

  • Bachelor's degree in computer science, software engineering, data science, machine learning, or related discipline.
  • 5+ years of software engineering and site reliability engineering experience operating large-scale, distributed, cloud-native systems in production.
  • Experience defining and owning Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Service Level Agreements (SLAs); error budgets; incident command and on-call; building and operating production observability; and environment integrity and drift prevention.
  • 3+ years of experience with cloud-native engineering and cloud platform ownership on any of the cloud hyperscalers such as Azure, AWS, or GCP.
  • Prior experience operating AI/ML and agentic workloads in production, understanding their reliability failure modes, MLOps/LLMOps, and the AI control plane.
  • Experience with load and performance testing under simulated production traffic, chaos engineering, capacity planning, autoscaling, and cloud/AI cost engineering.
  • Understanding of Business Context Diagrams (BCD), sequence/activity/state/entity relationship/data flow diagrams, OOP/OOD, data structures, algorithms, and code instrumentations, and AI-augmented spec-driven development.
  • Experience using methodologies & tools such as XP, Lean, DevSecOps, SRE, ADO, GitHub, SonarQube, MLflow, and agentic AI frameworks (e.g. LangFuse, LangSmith, or equivalent multi-agent orchestration tools).
  • Must be a U.S. Citizen or Green Card holder.

Similar jobs