Data Engineer
Haystack · Atlanta, GA · Yesterday
RemoteRemoteInformation TechnologyFull-time
About the Role
Design, develop, and maintain scalable ETL pipelines using AWS Glue and Apache Spark (PySpark). Build and orchestrate data workflows using AWS Step Functions. Design and implement S3 Data Lake architectures following AWS best practices. Develop and deploy containerized applications on Amazon EKS using Kubernetes. Build event-driven data processing solutions using Amazon SQS, AWS Lambda, and KEDA for auto-scaling. Manage metadata using AWS Glue Data Catalog. Implement data validation and governance using AWS Glue Data Quality. Monitor applications and data pipelines using Amazon CloudWatch. Optimize ETL jobs, Spark workloads, and Kubernetes deployments for performance and scalability.
Requirements
- Strong experience in building scalable cloud-native data platforms and ETL pipelines on AWS
- Hands-on expertise with AWS Glue, Spark ETL, Step Functions, Amazon EKS, Kubernetes, Lambda, S3 Data Lake architectures, SQS, KEDA, Glue Data Catalog, Glue Data Quality, and CloudWatch
- Proficiency in Apache Spark (PySpark)
- Experience with containerized applications and Kubernetes
- Ability to implement data validation and governance
- Strong analytical and problem-solving skills
Benefits
- Opportunity to work with cutting-edge cloud technologies
- Contribute to scalable cloud-native data platforms
- Full-time engagement with a dynamic team
- Remote work flexibility