Jobs · Information Technology · Texas

Senior Data Engineer, Infrastructure Reliability

Amazon · Austin, TX · 3 wk ago
Information TechnologyFull-time

Help build the data foundation that keeps Amazon's fulfillment network running 24/7. Infrastructure Reliability is building an AI-powered platform that detects, diagnoses, and resolves incidents across thousands of sites globally, and none of it works without clean, reliable, well-modeled data.

About the role

You will own and evolve production data pipelines that power real-time incident detection and correlation, build the data lake and unified data models that unblock ML and autonomous resolution initiatives, and rescue historical telemetry data before it expires and becomes permanently unrecoverable. This is a high-ownership role with direct, visible impact on Amazon's global fulfillment operations, working closely with applied scientists who train models on the data you build.

Responsibilities

  • Design, build, and operate ETL pipelines that ingest, transform, and correlate incident and telemetry data from sources including DynamoDB, OpenSearch, CloudWatch, and internal incident management systems, publishing curated datasets to our data lake and Andes for consumption by dashboards, applications, and ML models.
  • Take ownership of production pipelines, including scheduled batch processing jobs and Glue-based ETL workflows, ensuring reliability, monitoring, and timely resolution of pipeline issues.
  • Build data retention and ingestion pipelines to preserve time-bound telemetry signals before source-system expiration windows, and design standardized ingestion frameworks that normalize data from many disparate sources into a common, science-consumable format.
  • Contribute to the design of a unified data lake architecture, including schema design, partitioning strategy, and access patterns, replacing fragmented, duplicated pipelines with a single source of truth.
  • Partner closely with applied scientists and engineers to grasp data requirements for ML model training and feature engineering.
  • Mentor other engineers on data engineering best practices, code quality, and pipeline design.

A day in the life

You might start your day investigating a pipeline failure alert, tracing it back through Glue job logs to a schema change upstream. Later, you're in a design discussion with an applied scientist about what shape of data would best enable a new model, translating that into a concrete schema. In the afternoon, you're heads-down building a new ingestion adapter or reviewing a teammate's pull request. Your work directly determines what data is available, and reliable, for the platform's detection and reasoning capabilities.

About the team

Infrastructure Reliability sits within Amazon's Robotics organization, building the platform that keeps fulfillment operations running no matter what breaks. We do not own any single domain; we build the data and orchestration layer that sees across all of them, identifying failures that cascade across team boundaries. We are a small, technically deep team building AI-powered detection and remediation capabilities at scale. We value ownership, rigor, and hands-on technical depth, and we move quickly from idea to production.

Requirements

  • 5+ years of data engineering experience
  • Experience with data modeling, warehousing and building ETL pipelines
  • Experience with SQL
  • Experience in at least one modern scripting or programming language, such as Python, Java, Scala, or NodeJS
  • Experience mentoring team members on best practices
  • Bachelor's degree in computer science, engineering, analytics, mathematics, statistics, IT or equivalent

Preferred Qualifications

  • Experience with big data technologies such as: Hadoop, Hive, Spark, EMR
  • Experience operating large data warehouses
  • Master's degree in computer science, engineering, analytics, mathematics, statistics, IT or equivalent

Benefits

  • Medical, Dental, and Vision Coverage
  • Maternity and Parental Leave Options
  • Paid Time Off (PTO)
  • 401(k) Plan

Pay

USA, TX, Austin - $154,600.00 - $209,100.00 USD annually. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location.

Similar jobs