Data Engineer
Capgemini · New York, NY · Yesterday
Information TechnologyFull-time
About the role
We are seeking a highly skilled GCP Data Engineer with strong Python expertise to design, build, and optimize scalable data solutions on Google Cloud Platform (GCP). The ideal candidate will have hands-on experience developing batch and real-time data pipelines, working with large-scale datasets, and enabling analytics and AI/ML use cases.
Responsibilities
- Data Engineering & Pipeline Development: Develop and optimize ETL/ELT workflows for structured and unstructured data processing using GCP services such as Dataflow, Dataproc, and Pub/Sub; implement event-driven data processing using Cloud Functions and Pub/Sub; build and manage data ingestion frameworks for streaming and batch data sources.
- Data Storage & Processing: Design and optimize data lakes and data warehouses using BigQuery and Cloud Storage; develop efficient data models to support analytics, reporting, and machine learning workloads; optimize performance and cost of data pipelines and queries.
- Development & Automation: Develop solutions using Python; automate workflows and orchestration using Cloud Composer (Airflow); implement CI/CD pipelines and deployment automation.
- Collaboration & Support: Collaborate with analytics, AI/ML, and business teams for data consumption needs; troubleshoot data issues and perform root cause analysis; continuously improve pipeline reliability, scalability, and performance.
Required Technical Skills
- GCP Services: BigQuery, Dataflow, Dataproc, Pub/Sub, Cloud Functions (Gen2), Cloud Composer (Airflow), Cloud Storage, Cloud SQL
- Programming: Python (advanced)
- Query Language: SQL (advanced)
- Data Processing: Batch & Streaming architectures
- Scripting: Bash/Shell scripting
- Concepts: Data Warehousing, ETL/ELT, Data Lake / Lakehouse architectures
Preferred Skills
- Insurance Domain Experience | Certification
- Google Cloud Platform (GCP) services: Vertex AI, Cloud Run, BigQuery, Cloud Storage, IAM, Cloud Logging, Cloud Monitoring
- AI & Machine Learning: Large Language Models (LLMs), Gemini Models, Agentic AI Frameworks, Prompt Engineering, RAG Architectures, Vector Search Concepts, Semantic Retrieval, AI Evaluation Frameworks
- Data & Analytics: SQL, BigQuery, Metadata Modeling, Data Pipelines
Required Experience
- 5+ years of overall data engineering or software engineering experience
- 2+ years of hands-on Google Cloud Platform experience
- 2+ years of Python development
- 2+ years of experience building data pipelines (batch and streaming)
Preferred Qualifications
- Experience with Dataproc (Spark/PySpark) for large-scale processing
- Familiarity with event-driven architectures
- Knowledge of Terraform or Infrastructure as Code
- Understanding of cost optimization (FinOps)
- Google Cloud Professional Data Engineer Certification
- Experience supporting AI/ML data pipelines
Pay
The base compensation range for this role in the posted location is: $75,000–$96,000. Actual compensation may vary based on geographic location, education, certifications, relevant experience, seniority, market conditions, internal pay equity, and other legally permissible factors.