Big Data Hadoop Engineer
Staffing Spot, Inc. · Irving, TX · 5 days ago
HybridInformation TechnologyContract
About the role
We are seeking an experienced Senior Big Data Engineer to design and implement large-scale data solutions using PySpark, Hadoop, GCP, and graph databases.
Responsibilities
- Develop and maintain distributed data-processing applications with PySpark.
- Build and optimize data pipelines for both batch and streaming ingestion using GCP-native services.
- Design and implement data models and large-scale data processing solutions within the Hadoop ecosystem.
- Work with Neo4j and graph databases to model and query complex relationships.
- Implement CI/CD pipelines for data engineering workflows.
- Apply Google Cloud architecture best practices, including Cloud Storage bucket design, naming standards, lifecycle management, and IAM access-control mechanisms.
- Create and manage BigQuery views, APIs, and curated analytical datasets for data consumption and exposure.
Requirements
- 6+ years of experience in Data Engineering, including data pipelines, data modeling, and large-scale data processing.
- 4+ years of hands-on PySpark experience developing distributed data-processing applications.
- 4+ years of experience with the Hadoop ecosystem and HDFS.
- 4+ years of experience with GCP, including:
- Google Cloud architecture
- Cloud Storage bucket design and structuring
- Naming standards and lifecycle management policies
- IAM and access-control mechanisms
- Hands-on experience with Neo4j and graph databases.
- Strong experience implementing CI/CD pipelines.
Skills
- Hadoop ecosystem
- Data modeling
- PySpark
- GCP (Google Cloud Platform)
- Neo4j
- Data pipelines
- Big Data