Senior Director – Real World Data (RWD) Architect - Engineer
Eli Lilly and Company · Indianapolis, IN · 3 days ago
Engineering$170k–$249k/yrFull-time
Responsibilities
- Lead the design, development, and implementation of cloud-native data products and high-throughput data pipelines that transform raw real-world data into scalable, reliable, analysis-ready assets supporting analytics, reporting, and evidence generation.
- Lead the Analytic Data Products Strategy to deliver key data assets that enable streamlined, compliant execution and analytics.
- Own the end-to-to-end lifecycle of RWD data products, from requirements gathering and prototyping through production deployment and optimization, ensuring scalability, reliability, performance, and reproducibility across cloud environments (e.g., Databricks, AWS S3, Azure Data Lake).
- Build, optimize, and maintain ETL/ELT ingestion and transformation pipelines for large-scale, multi-modal RWD — including claims, complex EHR data, and other linked healthcare datasets — handling data volumes ranging from tens of millions to billions of records.
- Implement and manage lakehouse-style data architectures (e.g., medallion bronze/silver/gold patterns) using Databricks and cloud object storage (AWS S3, ADLS) to produce versioned, partitioned, and audit-ready data assets.
- Write and maintain reusable, version-controlled transformation logic incorporating healthcare coding and terminology standards (e.g., ICD-10/ICD-9, NDC, RxNorm, SNOMED, CPT/HCPCS, LOINC) to produce domain-level datasets such as demographics, diagnoses, treatments, procedures, encounters, and labs.
- Optimize SQL and distributed processing workloads (e.g., Spark-based jobs) for performance across very large datasets, applying partitioning, indexing, predicate pushdown, denormalization, and other optimization strategies appropriate to analytical workloads.
- Translate analytic, business, and research requirements into reproducible data extraction and transformation logic, supporting cohort construction, temporal logic, and consistent reuse of RWD across teams.
- Apply deep understanding of healthcare data structures and standards when engineering data products, ensuring datasets are fit for purpose for downstream analytics and compliant with scientific, regulatory, and audit expectations.
- Establish and implement standard engineering practices and methodology across the data asset lifecycle, including automated data ingestion, data quality checks, integrity testing, validation, monitoring, alerting, and documentation from source table to analysis-ready output.
- Lead CI/CD pipeline setup, code review, and testing standards, ensuring all transformation code is version-controlled, tested, and deployable in a reproducible manner.
- Collaborate closely with multi-functional partners — data scientists, statisticians, analytics leaders, and other technical teams — to understand business and technical requirements and develop documentation of RWD engineering standards, transformation templates, code list repositories, and pipeline performance guidelines.
- Provide technical consultation to collaborators on appropriate use of data products and underlying RWD assets, including structural limitations of specific data sources, join strategies, and performance considerations; develop source-specific training materials for HEOR scientists, SDIA, and statisticians.
- Develop and implement KPIs to measure system performance, efficiency and pull through to program impact.
- Create an inclusive culture where producing and maintaining high-quality data is a core discipline.
Requirements
- Bachelor’s degree in Computer Science, Engineering, Statistics, Information Technology, Bioinformatics or Technical Field.
- Minimum of 5 years of hands-on data engineering experience.
- Minimum of 5 years of applied expertise across Python, SQL, Java, Spring, Spring Boot, Prefect, and/or other business intelligence tools, ETL/ELT pipelines, and cloud platforms (AWS Glue/EMR, Snowflake, or Databricks), — applied directly to real-world healthcare data at scale.
- Minimum of 3 years of direct people management experience, including leading, coaching, and developing team members.
Qualifications
- Master’s degree in Computer Science, Engineering, Statistics, Information Technology, Bioinformatics.
- Experience with distributed computing frameworks (Spark, Dask) for large-scale RWD processing.
- Advanced SQL optimization skills for AWS Redshift and/or S3/Databricks-based architectures, including query tuning and workload management.
- Deep understanding of healthcare coding standards (ICD-10, NDC, RxNorm, SNOMED CT, CPT, LOINC).
Skills
- R fluency a plus for collaboration with statistical and HEOR teams on dataset validation and specification.
- Healthcare RWD — reviews and EHR Medical and pharmacy claims: deep working knowledge of CCAE, Optum Clinformatics, IQVIA PharMetrics, Truveta, and Komodo data structures — enrollment/eligibility tables, revenue codes, place-of-service codes, inpatient vs. outpatient claim splitting, drug identification at NDC and GPI level, and known structural quirks of each source.
- Complex EHR data: HL7 FHIR resources, Epic/Cerner/Truveta data models, clinical note schemas, problem list hierarchies, medication order and administration tables, lab result normalization, and vital sign time series — including semi-structured and nested JSON/XML from EHR exports and FHIR APIs.
- Healthcare terminology and ontology mapping: ICD-10-CM/PCS, ICD-9-CM, NDC, RxNorm, SNOMED CT, CPT-4, HCPCS, LOINC, ATC — building, versioning, and governing code list repositories joined to raw tables to produce concept-labeled analysis-ready datasets.
- OMOP CDM: transforming source claims and EHR data to OMOP v5.x including vocabulary loading (Athena), ETL specification documentation, and Achilles/DQD data quality checks.
- Phenotyping and cohort construction: translating clinical study protocols into reproducible extraction logic — index date derivation, washout periods, time-varying covariates, censoring — suitable for pharmacoepidemiology and HEOR studies.
- Regulatory and privacy frameworks: HIPAA-compliant data handling, de-identification standards (Safe Harbor, Expert Determination), DUA compliance, and audit trail requirements for FDA-grade RWE submissions.
- AI Fluency Deploy NLP pipelines for structured extraction from unstructured clinical notes — operationalizing pre-trained biomedical language models (BioBERT, ClinicalBERT) for outcome ascertainment and phenotyping within the data pipeline.
- Implement AI-powered data profiling and anomaly detection to automate quality checks across large ingestion runs and surface issues before they reach analytical teams.
- Use LLM-assisted SQL generation and code review tools to accelerate pipeline development and reduce query errors at scale.
- Apply intelligent caching, query result reuse, and automated feature engineering to optimize compute costs and prepare multi-modal variables for downstream ML inputs.
- Familiarity with MLOps tooling (MLflow, SageMaker, Azure ML) for versioning and monitoring AI models integrated with data pipelines.