Mid-Level Data Engineer
Largeton Group · Columbus, OH · 1 mo ago
On-siteInformation TechnologyContract
About the role
Support the IAM Data Lake Engineering initiative. Build and maintain Data Lake solutions on Google Cloud Platform (GCP) using big data tools and technologies.
Responsibilities
- Design, develop, and optimize data pipelines for ingestion, processing, and transformation of large-scale datasets.
- Implement data modeling and data processing solutions to support analytical and operational requirements.
- Utilize PySpark for distributed data processing and transformation tasks.
- Work extensively with the Hadoop ecosystem, including HDFS for data storage and management.
- Apply deep knowledge of GCP architecture, including bucket structuring, naming conventions, lifecycle policies, and access controls.
- Integrate and manage data using Neo4j for graph database requirements.
- Ensure robust CI/CD processes for continuous integration and deployment of data engineering solutions.
- Build both batch and streaming data ingestion pipelines leveraging GCP-native services.
- Develop and maintain data consumption and exposure layers via views, APIs, and curated analytical datasets.
- Collaborate with cross-functional teams to ensure data solutions are scalable, secure, and aligned with organizational standards.
Schedule
Duration of assignment is 12 months, with a possible extension based on project needs and performance.