Senior Product Manger - Tech, Infrastructure Reliability
About the role
Join Amazon's Fulfillment Technologies & Robotics (FTR) team to spearhead the product vision for a platform that ensures Amazon's fulfillment network never stops — even as we move toward fully self-governing, zero-touch operations. You'll own the roadmap for an AI-powered infrastructure reliability platform that prevents, detects, and resolves incidents across thousands of fulfillment sites globally. This is a rare opportunity for a technically deep product leader who can write code, deliver proof-of-concepts, and engage as a peer with data scientists and engineers. You will shape how LLMs, multi-agent systems, and machine learning are applied to one of the most operationally critical platforms Amazon has ever built — and your hands-on technical contributions will directly accelerate the team's ability to move from idea to production.
Responsibilities
- Own and drive the multi-year product roadmap for the Infrastructure Reliability AI-Ops platform, spanning three strategic programs: zero-touch incident resolution, associate-directed work tooling, and predictive failure prevention. Define the vision, strategy, and success metrics for AI-powered progressive detection, incident consolidation, self-governing remediation orchestration, and cross-domain observability capabilities that serve thousands of fulfillment sites globally.
- Go beyond traditional product management by writing code and delivering working proof-of-concepts that validate technical hypotheses before committing engineering resources. Prototype multi-agent reasoning pipelines, explore new anomaly detection approaches, or stress-test LLM prompt chains against real incident data to compress the distance between idea and validated direction.
- Apply deep knowledge of machine learning fundamentals to shape how the platform detects, consolidates, and reasons about failures. Engage meaningfully with data scientists on model architecture selections, feature engineering tradeoffs, and evaluation frameworks — understanding not just what a model produces but why, and whether that reasoning can be trusted in production.
- Define how AI reasoning techniques — including chain-of-thought prompting, retrieval-augmented generation, confidence calibration, and evidence accumulation — build progressive confidence about incident severity and failure origin rather than making binary selections from rigid thresholds. Shape how LLMs are applied to diagnostic summarization, resolution suggestion, and automated stakeholder communication.
- Design the multi-agent architecture that orchestrates detection, investigation, consolidation, diagnosis, and remediation as a coordinated system. Work with engineering to define agent roles, communication protocols, handoff conditions, and safety boundaries to ensure self-governing agents act with appropriate confidence and escalate when uncertainty is high.
- Translate complex operational and technical requirements into a prioritized backlog, making clear tradeoffs between feature depth, platform scalability, and autonomous site readiness milestones. Serve as the voice of Incident Managers, domain engineers, and Operations Control Center stakeholders, advocating for their needs during executive-level planning.
- Define and track the business case across all three programs — including mean time to resolve improvements, lost labor hour reduction, and first page resolution improvement — to secure continued investment. Establish mechanisms to measure platform performance against key metrics like auto-detection rate, false positive rate, consolidation accuracy, and remediation success rate.
- Drive cross-functional alignment across Fulfillment Technologies, Robotics, Network Engineering, Application teams, and Operations to ensure the platform's cross-domain orchestration model is well understood and adopted. Lead executive-level reviews of program progress, risks, and investment cases.
A day in the life
You spend most of your time at the intersection of product strategy and hands-on technical work. A typical day might start by pulling incident data into a notebook to test a new detection signal, then jumping into a whiteboard session with engineers debating multi-agent handoff reasoning. You might prototype a diagnostic flow in the afternoon just to prove a concept is worth building. Occasionally, you will find yourself in the operations center watching real operators work through a network failure — because staying grounded in how people actually experience the platform is what separates good product selections from great ones.
Requirements
Basic Qualifications
- Bachelor's degree
- Experience owning/driving roadmap strategy and definition
- Experience with feature delivery and tradeoffs of a product
- Experience contributing to engineering discussions around technology decisions and strategy related to a product
- Experience managing technical products or online services
- Experience in representing and advocating for a variety of critical customers and stakeholders during executive-level prioritization and planning
Preferred Qualifications
- Experience in using analytical tools, such as Tableau, Qlikview, or QuickSight
- Experience in building and driving adoption of new tools
About the team
The Infrastructure Reliability team sits within Amazon's Robotics organization, operating as the cross-domain orchestration layer for a fulfillment network that processes customer orders continuously across thousands of sites. Our mission is simple and purposeful: operations never stop, no matter what breaks. We do not own any single domain — instead, we build the platform that sees across all of them, identifying failures that cascade across team boundaries and coordinating the capabilities that domain teams have built to resolve those failures faster than any single team could alone.
We are now building the AI-powered platform that applies machine learning, reasoning, and multi-agent orchestration to take our results from promising to industry-defining. We value expert rigor, customer obsession, and hands-on technical depth. The ideal teammate is as comfortable writing a proof-of-concept as they are writing a product strategy document. If you want to work on a problem that is technically fascinating, operationally critical, and commercially enormous, this is the team for you.
Benefits
- Medical, Dental, and Vision Coverage
- Maternity and Parental Leave Options
- Paid Time Off (PTO)
- 401(k) Plan
- Sign-on payments and restricted stock units (RSUs)
- Health insurance (including medical, dental, vision, prescription, Basic Life & AD&D insurance, and optional supplemental life plans)
- Employee Assistance Program (EAP) and Mental Health Support
- Medical Advice Line
- Flexible Spending Accounts
- Adoption and Surrogacy Reimbursement coverage
- 401(k) matching
- Parental leave
Pay
The base salary range for this position varies by location:
- North Reading, MA: $151,200 - $204,600 USD annually
- Nashville, TN: $143,700 - $194,300 USD annually
- Austin, TX: $151,200 - $204,600 USD annually
- Arlington, VA: $151,200 - $204,600 USD annually
Final compensation will be determined based on factors including experience, qualifications, and location.