Jobs · Project Management · Massachusetts

Senior Product Manger - Tech, Infrastructure Reliability

Amazon · Boston, MA · Today
Project ManagementFull-time

Help build the AI-powered platform that keeps Amazon's fulfillment network running 24/7 — even as we move toward fully self-governing, zero-touch operations. This is a rare hands-on Product Manager role where you'll write code, build proof-of-concepts, and shape how LLMs and multi-agent systems detect, diagnose, and resolve infrastructure incidents across thousands of sites globally.

About the role

You will own the multi-year roadmap for zero-touch incident resolution, associate-directed tooling, and predictive failure prevention, while prototyping and validating technical concepts directly — from anomaly detection approaches to multi-agent reasoning pipelines. You'll partner daily with data scientists and engineers on model architecture, LLM reasoning techniques, and multi-agent orchestration, influencing one of the most technically ambitious and operationally critical platforms in Amazon's Robotics organization. The team is small, the problem is large, and the right candidate will have significant ownership over both technical direction and product strategy from day one.

Responsibilities

  • Own and drive the multi-year product roadmap for the Infrastructure Reliability AI-Ops platform across three strategic programs: zero-touch incident resolution, associate-directed work tooling, and predictive failure prevention.
  • Write code and deliver working proof-of-concepts — prototyping multi-agent reasoning pipelines, testing anomaly detection approaches, and stress-testing LLM prompt chains against real incident data — to validate technical hypotheses before committing engineering resources.
  • Apply deep machine learning knowledge to shape how the platform detects, consolidates, and reasons about failures, engaging meaningfully with data scientists on model architecture selections, feature engineering tradeoffs, and evaluation frameworks.
  • Define how the platform uses chain-of-thought prompting, retrieval-augmented generation, and confidence calibration to build progressive confidence about incident severity and origin.
  • Define the multi-agent architecture coordinating detection, investigation, diagnosis, and remediation, including agent roles, handoff conditions, and safety boundaries.
  • Translate operational needs into a prioritized backlog, representing Incident Managers, domain engineers, and Operations Control Center stakeholders during executive-level planning, while defining and tracking the business case and driving cross-functional alignment across Fulfillment Technologies, Robotics, and Network Engineering.

A day in the life

You spend your time at the intersection of product strategy and hands-on technical work. A typical day might start by pulling incident data into a notebook to test a new detection signal, then jumping into a whiteboard session with engineers debating multi-agent handoff logic. You might prototype a diagnostic flow in the afternoon just to prove a concept is worth building. Occasionally you will find yourself in the operations center watching real operators work through a network failure, because staying grounded in how people actually experience the platform is what separates good product selections from great ones.

About the team

The Infrastructure Reliability team sits within Amazon's Robotics organization, operating as the cross-domain orchestration layer for a fulfillment network that processes customer orders continuously across thousands of sites. Our mission is simple and purposeful: operations never stop, no matter what breaks. We do not own any single domain — instead, we build the platform that sees across all of them, identifying failures that cascade across team boundaries and coordinating capabilities faster than any single team could alone. We are now building the AI-powered platform that applies machine learning, reasoning, and multi-agent orchestration to take our results from promising to industry-defining.

Requirements

  • 5+ years of technical product management with internet business experience
  • 5+ years of working as a Technical Product Manager
  • 3+ years of technical (software development, network development, IT, or other related) experience
  • 7+ years of full product life cycle experience
  • 5+ years of creating written docs for development of new products
  • 5+ years of enterprise security product experience
  • 5+ years of product management in the cloud computing technology space
  • Bachelor's degree in computer science, engineering, math, finance, or economics
  • Experience in taking a product from conception & definition phase through engineering design and taking it to market
  • Experience delivering large-scale SaaS, PaaS or IaaS products where you are responsible for the full product lifecycle, from concept through GTM (go to market)

Preferred Qualifications

  • Knowledge of engineering practices and patterns for the full software/hardware/networks development life cycle, including coding standards, code reviews, source control management, build processes, testing, certification, and livesite operations
  • Experience working within teams delivering software products and features using agile methodologies

Benefits

  • Medical, Dental, and Vision Coverage
  • Maternity and Parental Leave Options
  • Paid Time Off (PTO)
  • 401(k) Plan

Pay

  • USA, MA, North Reading: $152,200 - $205,900 USD annually
  • USA, TN, Nashville: $144,600 - $195,600 USD annually
  • USA, TX, Austin: $152,200 - $205,900 USD annually
  • USA, VA, Arlington: $152,200 - $205,900 USD annually

Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave.

Similar jobs