Data Engineer II, Business Data Technologies
About the role
Amazon’s eCommerce Foundation (eCF) organization is responsible for the core components that drive the Amazon website and customer experience. Serving millions of customer page views and orders per day, eCF builds for scale. As an organization within eCF, the Business Data Technologies (BDT) group is no exception. We collect petabytes of data from thousands of data sources inside and outside Amazon including the Amazon catalog system, inventory system, customer order system, page views on the website and Alexa systems. We also support Amazon subsidiaries such as IMDB and Audible. We provide interfaces for our internal customers to access and query the data hundreds of thousands of times per day, using Amazon Web Service’s (AWS) Redshift, Hive, and Spark. We build scalable solutions that grow with the Amazon business.
Responsibilities
- Work in one of the world's largest cloud-based data lakes, designing, developing, implementing, testing, and operating large-scale, high-volume, high-performance data structures for analytics and deep learning.
- Implement data ingestion routines both real time and batch using best practices in data modeling, ETL/ELT processes leveraging AWS technologies and Big data tools.
- Gather business and functional requirements and translate these requirements into robust, scalable, operable solutions that work well within the overall data architecture.
- Analyze source data systems and drive best practices in source teams.
- Participate in the full development life cycle, end-to-end, from design, implementation and testing, to documentation, delivery, support, and maintenance.
- Produce comprehensive, usable dataset documentation and metadata.
- Evaluate and make decisions around dataset implementations designed and proposed by peer data engineers.
- Evaluate and make decisions around the use of new or existing software products and tools.
- Mentor junior data engineers.
Qualifications
- 3+ years of data engineering experience.
- Experience with data modeling, warehousing and building ETL pipelines.
- Preferred: Experience with AWS technologies like Redshift, S3, AWS Glue, EMR, Kinesis, FireHose, Lambda, and IAM roles and permissions.
- Preferred: Experience with non-relational databases / data stores (object storage, document or key-value stores, graph databases, column-family databases).