Senior Data Engineer
CGI · Pittsburgh, PA · 2 days ago
Information Technology$71k/yrFull-time
This role is based at our client site in Pittsburgh, PA and requires on-site presence 5 days a week.
About the role
We are seeking a Data Engineer with 6+ years of experience to design and maintain scalable data pipelines supporting analytics, reporting, and operational needs. The role involves collaborating with cross-functional teams to ensure data alignment with business requirements and enterprise standards.
Responsibilities
- Design and build scalable data pipelines aligned with business needs
- Process large datasets (batch and near real-time)
- Ensure data quality, consistency, and governance standards across systems
- Support data integration and transformation efforts for analytics and reporting platforms
- Maintain data dictionaries, metadata, and documentation
- Participate in data architecture reviews and model validation processes
- Support analytics reporting and risk platforms
Requirements
- Bachelor’s degree in Computer Science or related field
- 6+ years of experience in data engineering and big data processing
- 3+ years of experience building Spark Streaming and Kafka streaming
- Strong expertise in Apache Spark (Spark Core, Spark SQL) and PySpark for large-scale batch processing
- Strong expertise in Neo4J (Native Graph Database Management system)
- Experience working with structured and semi-structured data, including complex transformations and performance tuning
- Proficiency in data ingestion and integration from sources like Oracle, SQL Server, Hive, HDFS, and S3; transform data into curated data models
- Experience writing data to Hive tables, Data Lakes (Iceberg), and downstream reporting systems
- Strong knowledge of SQL and data modeling concepts
- Hands-on experience with Apache Airflow for workflow orchestration (DAG design, scheduling, monitoring)
- Proficiency in shell scripting for job automation, file validation, dependency handling, and logging (e.g., triggering Spark jobs, file checks, archiving, purging, managing job dependencies)
- Strong understanding of batch processing and batch job scheduling frameworks
- Experience migrating from CA7/Control-M to Airflow (daily, hourly, weekly schedules)
- CI/CD for data pipelines
- Experience ensuring data quality, reliability, and compliance in regulated environments
- Good communication and documentation skills
Pay
A reasonable estimate of the compensation range for this role in the U.S. is $79,600.00 – $139,300.00.
Benefits
- Competitive compensation
- Comprehensive insurance options
- Matching contributions through the 401(k) plan and the share purchase plan
- Paid time off for vacation, holidays, and sick time
- Paid parental leave
- Learning opportunities and tuition assistance
- Wellness and well-being programs