Senior Data Engineer
CGI · Strongsville, OH · 2 wk ago
Information Technology$80k/yrFull-time
About the role
The successful candidate will serve as a technical leader on a large scale digital transformation initiative, collaborating with cross-functional business and technology teams while guiding both onshore and offshore development teams to deliver high-quality, scalable data solutions.
Responsibilities
- Design, develop, and maintain scalable data pipelines and transformation frameworks using Hadoop ecosystem technologies.
- Build high-performance data transformation solutions utilizing Hive, Spark, Python, Impala, and Apache Iceberg.
- Develop robust ETL/ELT processes supporting enterprise-scale analytics and data products.
- Design and optimize Hive, Spark SQL, and Impala queries for large datasets.
- Implement scalable data lake solutions leveraging Apache Iceberg table formats and best practices.
- Collaborate closely with Product Owners, Business Analysts, Subject Matter Experts, Architects, and Technical Managers to translate business requirements into technical solutions.
- Lead technical design discussions and establish engineering best practices across the development team.
- Provide technical leadership and mentoring to both onshore and offshore development teams.
- Conduct code reviews and ensure adherence to coding standards, performance optimization, and data quality best practices.
- Troubleshoot production issues and implement long-term sustainable solutions.
- Drive continuous improvement through automation, reusable frameworks, and engineering best practices.
- Participate in Agile ceremonies including sprint planning, backlog refinement, and retrospectives.
- Stay current with emerging Big Data, AWS cloud, and AI technologies and recommend innovative solutions.
Requirements
- 6+ years of experience in Data Engineering, Big Data development, or enterprise data platforms.
- Strong experience developing enterprise-scale Big Data applications.
- Hands-on expertise with: Hadoop ecosystem, Hive, Apache Spark, Impala, PySpark, Apache Iceberg, and strong SQL development and query optimization skills.
- Experience designing scalable ETL/ELT pipelines.
- Experience working with distributed computing environments and large-volume datasets.
- Strong analytical, troubleshooting, and problem-solving skills.
- Experience working in Agile/Scrum delivery environments.
- Excellent communication and stakeholder collaboration skills.
- Proven ability to lead technical initiatives across geographically distributed teams.
- Experience with cloud-based data platforms (AWS, Azure, or Google Cloud).
- Experience with orchestration tools such as Airflow or Oozie.
- Experience with version control systems such as Git and CI/CD pipelines.
- Familiarity with data governance, metadata management, and data quality frameworks.
- Experience with performance tuning and optimization of Big Data workloads.
- Working knowledge of Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) architectures is a plus.
Qualifications
- To be successful in this role, candidates should have a bachelor's degree in Computer Science, Information Technology, or a related field.
- Experience with Generative AI technologies and AI-assisted development is highly desirable.
Benefits
This position can be performed onsite five days a week at our client site in Strongsville, OH or Pittsburgh, PA or Dallas, TX.
Pay
$79,600.00 - $139,300.00 per year
Schedule
Full-time, 40 hours per week