Data Engineer
About the role
You’ll make a difference by designing, building, and maintaining scalable data products and data infrastructure on AWS. You will develop robust data pipelines using AWS-native services and integrate data across multiple application silos. You will support the automation, deployment, and operation of AI/ML workflows. You will define data ingestion strategies, data formats and schemas, metadata/catalog integrations, federated data access, and end-to-end data product creation. You will enable advanced analytics, AI/ML, and data-driven use cases by designing efficient data-access patterns and tooling (including support for LLM context engineering). You will collaborate closely with Data Scientists, ML Engineers, and Product teams to translate business needs into scalable data solutions. You will actively participate in design discussions, technical reviews, and cross-team collaboration forums.
Responsibilities
- Design, build, and maintain scalable data products and data infrastructure on AWS.
- Develop robust data pipelines using AWS-native services and integrate data across multiple application silos.
- Support the automation, deployment, and operation of AI/ML workflows.
- Define data ingestion strategies, data formats and schemas, metadata/catalog integrations, federated data access, and end-to-end data product creation.
- Enable advanced analytics, AI/ML, and data-driven use cases by designing efficient data-access patterns and tooling (including support for LLM context engineering).
- Collaborate closely with Data Scientists, ML Engineers, and Product teams to translate business needs into scalable data solutions.
- Participate in design discussions, technical reviews, and cross-team collaboration forums.
Requirements
- Hold a Bachelor’s or Master’s degree in Computer Science, Engineering, or a related discipline.
- Have 3+ years of hands-on experience working with data at scale.
- Strong programming skills in Python and SQL.
- Strong understanding of databases (SQL and NoSQL), complex SQL queries and their performance across large, distributed databases.
- Experience in building and maintaining data pipelines.
- Proven experience in building and maintaining CI/CD pipelines.
- Solid understanding of SQL and NoSQL databases, complex SQL queries, and performance optimization across large, distributed systems.
- Practical experience with Apache Spark (preferably PySpark).
- Clear communication skills and the ability to capture and define technical requirements effectively.
Qualifications
- Experience with TypeScript and/or JavaScript.
- Working knowledge of databases such as PostgreSQL, DynamoDB, or similar technologies.
- Strong collaboration experience with Data Scientists and Machine Learning Engineers.
- Interest or hands-on exposure to MLOps and the AI product lifecycle.
- Domain experience (or strong willingness to learn) in IoT, time-series data, automation systems, digital twins, or smart buildings technologies.
- Experience with Infrastructure as Code (ideally CDK, CloudFormation, Terraform) and cloud-native technologies in AWS (e.g. Lambda, Athena, Glue, SageMaker, S3, etc.).
- Experience with tools like EMR, Snowflake, AWS Glue.
Benefits
- Flexible and hybrid working opportunities.
- A diverse, inclusive, and collaborative culture.
- Continuous learning and development opportunities.
- An attractive and competitive compensation package.