Senior Data Engineer
BioSpace · Indianapolis, IN · Yesterday
Information Technology$132k–$244k/yrFull-time
Job Responsibilities
- Lead the design, development, and implementation of complex data pipelines using various ETL/ELT tools and programming languages (e.g., Python, Scala, SQL).
- Develop and optimize data models, schemas, and data warehousing solutions (e.g., Databricks, AWS Redshift) to ensure high performance and data integrity.
- Serve as a recognized authority across multiple data technologies and business domains, independently resolving ambiguous or complex data problems.
- Collaborate with data scientists, analysts, and other collaborators to understand data requirements and translate them into technical solutions.
- Implement and maintain data quality checks, data governance policies, and data security measures.
- Troubleshoot and resolve data pipeline, infrastructure, and platform issues, ensuring data availability and reliability.
- Evaluate and recommend new data technologies and tools to improve efficiency and capabilities.
- Challenge existing approaches and recommend improvements to team and business data processes that drive innovation.
- Mentor and coach data engineers by sharing knowledge in specialized technologies and business domains to increase team technical growth.
- Promote ideas and influence technical decisions across multiple teams, including business and non-engineering collaborators.
- Participate in code reviews and ensure adherence to coding standards and architectural guidelines.
- Document data flows, processes, and system architectures comprehensively.
Qualifications
- Bachelor's degree in computer science, Engineering, Information Systems, or a related quantitative field.
- 8+ years of experience in data engineering, data warehousing, or a related field.
- Proven expertise in designing and building scalable data pipelines using cloud-based platforms (e.g., AWS, Azure, GCP).
- Strong hands-on experience with Databricks (e.g., Delta Lake, notebooks, workflows, Unity Catalog) for building and optimizing large-scale data solutions.
- Hands-on experience with Apache Airflow for orchestrating, scheduling, and monitoring complex data pipelines.
- Strong proficiency in SQL and at least one programming language (Python, Scala, Java).
- Extensive experience with data warehousing concepts and technologies (e.g., Airflow, Databricks, Redshift, Teradata).
- Experience with data visualization tools (e.g., Tableau, Power BI, Looker) to support analytics, reporting, and stakeholder communication.
- Solid understanding of data modeling techniques (dimensional, relational, star/snowflake schemas).