Kafka/Spark Developer
CGI · Pittsburgh, PA · 1 mo ago
OTHR$71k–$157k/yrFull-time
About the Role
CGI is seeking mid-level Kafka and Spark Software Developers to join our Applications Development and Maintenance team, supporting a large US Bank in an advanced technology environment. This role requires on-site presence 5 days a week in Pittsburgh, PA.
Responsibilities
- Develop and maintain scalable big data solutions using Hadoop, Spark, Kafka, and Impala to support enterprise data processing and analytics initiatives.
- Design, build, and optimize batch and real-time data pipelines for ingesting, processing, transforming, and delivering large volumes of structured and unstructured data.
- Develop Spark applications using PySpark, Scala, or Java for data transformation, aggregation, cleansing, and analytical processing.
- Build and maintain Kafka producers, consumers, topics, and streaming workflows to enable reliable real-time data ingestion and event-driven architectures.
- Design and implement logical and physical data models to support data warehousing, reporting, analytics, and business intelligence requirements.
- Monitor, troubleshoot, and tune Kafka and Spark streaming jobs to improve performance, scalability, and operational reliability.
- Optimize Hadoop ecosystem components, Spark jobs, Kafka configurations, and Impala queries to improve system performance and resource utilization.
- Collaborate with architects, data engineers, DevOps teams, and business stakeholders to design and implement modern streaming and event-driven data platforms.
- Analyze user requirements and define technical project scope and assumptions for assigned tasks.
- Create technical designs for new systems and modifications to existing systems.
- Translate detailed requirements into functional system designs.
- Prioritize work, meet deadlines, and establish effective working relationships with clients, project team members, supervisors, and employees from other departments.
- Partner with business leaders, enterprise architects, and product owners to identify new graph-based use cases, evaluate emerging technologies, and align initiatives with digital transformation goals.
Requirements
- At least 5+ years of experience in Big Data development, data engineering, or distributed data processing environments.
- Strong hands-on experience with Apache Kafka, including topic configuration, producer/consumer development, Kafka Connect, and Schema Registry.
- Extensive experience developing real-time data processing applications using Apache Spark Streaming and/or Spark Structured Streaming.
- Proficiency in Java, Scala, or Python (PySpark) with strong object-oriented programming and software development skills.
- Proficiency in writing and optimizing complex SQL queries using Impala, Hive, or similar distributed query engines.
- Hands-on experience with Hadoop ecosystem components including HDFS and Hive.
- Experience integrating Kafka and Spark with relational databases, NoSQL databases, cloud storage platforms, and enterprise applications.
- Strong analytical, troubleshooting, and performance tuning skills in distributed streaming environments.
- Excellent communication, collaboration, and stakeholder management skills, with the ability to work effectively in Agile/Scrum teams.
- Experience working in Agile development environments with strong collaboration, technical leadership, problem-solving, and stakeholder communication skills.
Pay
A reasonable estimate of the current compensation range for this role in the U.S. is $70,800.00 - $156,700.00.
Benefits
- Competitive compensation.
- Comprehensive insurance options.
- Matching contributions through the 401(k) plan and the share purchase plan.
- Paid time off for vacation, holidays, and sick time.
- Paid parental leave.
- Learning opportunities and tuition assistance.
- Wellness and well-being programs.