Associate Director of RWD Engineering
Eli Lilly and Company · Indianapolis, IN · 3 days ago
Engineering$156k–$229k/yrFull-time
Responsibilities
- Lead the design, development, and implementation of cloud-native data products and high-throughput data pipelines that transform raw real-world data into scalable, reliable, analysis-ready assets supporting analytics, reporting, and evidence generation.
- Lead the Analytic Data Products Strategy to deliver key data assets that enable streamlined, compliant execution and analytics.
- Own the end-to-end lifecycle of RWD data products, from requirements gathering and prototyping through production deployment and optimization, ensuring scalability, reliability, performance, and reproducibility across cloud environments (e.g., Databricks, AWS S3, Azure Data Lake).
- Build, optimize, and maintain ETL/ELT ingestion and transformation pipelines for large-scale, multi-modal RWD — including claims, complex EHR data, and other linked healthcare datasets — handling data volumes ranging from tens of millions to billions of records.
- Implement and manage lakehouse-style data architectures (e.g., medallion bronze/silver/gold patterns) using Databricks and cloud object storage (AWS S3, ADLS) to produce versioned, partitioned, and audit-ready data assets.
- Write and maintain reusable, version-controlled transformation logic incorporating healthcare coding and terminology standards (e.g., ICD-10/ICD-9, NDC, RxNorm, SNOMED, CPT/HCPCS, LOINC) to produce domain-level datasets such as demographics, diagnoses, treatments, procedures, encounters, and labs.
- Optimize SQL and distributed processing workloads (e.g., Spark-based jobs) for performance across very large datasets, applying partitioning, indexing, predicate pushdown, denormalization, and other optimization strategies appropriate to analytical workloads.
- Translate analytic, business, and research requirements into reproducible data extraction and transformation logic, supporting cohort construction, temporal logic, and consistent reuse of RWD across teams.
- Apply deep understanding of healthcare data structures and standards when engineering data products, ensuring datasets are fit for purpose for downstream analytics and compliant with scientific, regulatory, and audit expectations.
- Establish and implement standard engineering practices and methodology across the data asset lifecycle, including automated data ingestion, data quality checks, integrity testing, validation, monitoring, alerting, and documentation from source table to analysis-ready output.
- Contribute to CI/CD pipeline setup, code review, and testing standards, ensuring all transformation code is version-controlled, tested, and deployable in a reproducible manner.
- Collaborate closely with multi-functional partners — data scientists, statisticians, analytics leaders, and other technical teams — to understand business and technical requirements and develop documentation of RWD engineering standards, transformation templates, code list repositories, and pipeline performance guidelines.
- Provide technical consultation to collaborators on appropriate use of data products and underlying RWD assets, including structural limitations of specific data sources, join strategies, and performance considerations; collaborate with the RWD Operations Lead to develop source-specific training materials for HEOR scientists, SDIA, and statisticians.
- Create an inclusive culture where producing and maintaining high-quality data is a core discipline.
Requirements
- Bachelor’s degree in Computer Science, Engineering, Statistics, Information Technology, Bioinformatics, or Technical Field.
- Minimum of 3 years of hands-on data engineering experience with a demonstrated focus on healthcare or life sciences RWD.
- Minimum of 3 years of applied expertise across Python, SQL, Java, Spring, Spring Boot, Prefect, and/or other business intelligence tools, ETL/ELT pipelines, and cloud platforms (AWS Glue/EMR, Snowflake, or Databricks), — applied directly to real-world healthcare data at scale.
Additional Preferences
- Master’s degree in Computer Science, Engineering, Statistics, Information Technology, Bioinformatics.
- Experience with distributed computing frameworks (Spark, Dask) for large-scale RWD processing.
- Deep understanding of healthcare coding standards (ICD-10, NDC, RxNorm, SNOMED CT, CPT, LOINC), real-world data structures, and major RWD vendors and platforms (Truveta, Optum, IQVIA, Komodo, HealthVerity).
- Familiarity with DevOps and CI/CD practices relevant to data pipeline development and deployment.
- Strong problem-solving skills, attention to detail, and ability to work independently and collaboratively.