Senior Data Engineer with Pyspark
Infowave Systems, Inc · Rocky Hill, CT · 1 mo ago
On-siteInformation TechnologyFull-time
About the role
We are looking for an experienced and motivated Data Engineer with expertise in PySpark to join our dynamic team. As a key member of our data engineering team, you will play a crucial role in designing, building, and maintaining scalable data pipelines that enable efficient data processing and analytics within the healthcare domain.
Responsibilities
- Design, develop, and maintain scalable ETL pipelines using PySpark to process large datasets.
- Collaborate with cross-functional teams (data scientists, analysts, business stakeholders) to understand data requirements and deliver high-quality solutions.
- Work on administrative tasks, including monitoring, troubleshooting, and optimizing data pipelines and infrastructure.
- Manage data integration across healthcare systems, ensuring compliance with relevant standards.
- Leverage SSIS for ETL development and ensure smooth data movement across different environments.
- Integrate and transform data from multiple sources, ensuring data quality and consistency.
- Handle and resolve data processing issues, ensuring minimal disruption to operations.
- Document best practices, processes, and workflows to maintain pipeline efficiency and scalability.
- Work with both relational and non-relational databases, ensuring smooth data flow and optimized performance.
Requirements
- Strong experience in Data Engineering, with expertise in designing, building, and maintaining ETL pipelines.
- Strong proficiency in PySpark for large-scale data processing and transformation.
- Experience with ETL tools, particularly SSIS (SQL Server Integration Services).
- Solid understanding of data modeling, relational databases, and data warehousing principles.
- Experience working with cloud-based data storage and processing technologies (AWS, GCP, or Azure).
- Familiarity with healthcare data standards, such as HL7 and FHIR, is highly desirable.
- Proven ability in data pipeline monitoring, troubleshooting, and performance tuning.
- Strong communication skills and the ability to work collaboratively with cross-functional teams.