Jobs · Information Technology · Indiana

Senior CE Data Engineer

BioSpace · Indianapolis, IN · 2 days ago
Information Technology$63k–$150k/yrFull-time

About the role

At Lilly, the work is demanding because patients are waiting. We unite caring with discovery to help make life better for people around the world, knowing that every decision, every detail, and every day matters. Headquartered in Indianapolis, Indiana, our over 50,000 employees around the globe take on complex challenges to discover and deliver life-changing medicines, strengthen how health is understood and managed, and support the communities we serve. This is hard, urgent, selfless work—but it’s work worth doing. If you’re driven by purpose and ready to bring your best to work that truly matters for patients, we invite you to join us.

Lilly CIA powers the enterprise with governed, AI-ready data products, a converged semantic layer, and the identity and policy infrastructure behind Lilly's digital experiences. Within CIA, the LillyDirect Covered Entity (CE) Pharmacy pod builds and operates the data platform that sits inside the pharmacy's covered-entity trust boundary — handling identified PHI under HIPAA with strict isolation from parent-side systems. This role implements that platform on Databricks.

Responsibilities

  • Design and build scalable, efficient Databricks pipelines implementing the CE Architect's canonical data models across the medallion (bronze → silver → gold), scoped entirely within the CE trust boundary.
  • Evaluate and apply Databricks capabilities and integration patterns — Unity Catalog, Delta Lake, Databricks Workflows, serverless compute, Lakebase, ingestion connectors — selecting the right tool for each pipeline given performance, cost, and scalability constraints.
  • Implement the identity crosswalk and PHI classification logic defined by the CE Architect (Reltio MDM, Auth0/Passport, Datavant tokens, Scriptly, Genesys, Prescryptive, Transcend), safeguarding merge/split integrity in code.
  • Build to ODCS data contracts — implement schema, quality, freshness/SLA, and lineage requirements as enforced pipeline logic, not documentation.
  • Implement row/column-level security, masking, and tokenization boundaries so PHI isolation is enforced at the platform layer, in partnership with the Policy-as-Code Engineer's OPA/Rego policies.
  • Own end-to-end pipeline development lifecycle — requirements to prototyping to production deployment and maintenance — for CE data products, in partnership with Analytics Engineers and the DataHub Product Owner.
  • Build and maintain CI/CD pipelines (Github Actions [CI], Git-based promotion dev → test → prod) for CE data, contract, policy, and agent artifacts.
  • Automate data ingestion and product creation to reduce manual pipeline maintenance and onboarding time for new CE data sources.
  • Partner with the CE Data Architect on reference architecture and patterns, providing implementation feedback that keeps designs buildable and performant at scale.
  • Build the data pipelines that lets CE Skills and agents consume governed data — ensuring PHI classification and consent travel with the data into agentic consumption paths.

Requirements

  • Bachelor's degree in Computer Science, Engineering, Information Technology, or similar degree.
  • Experience in data engineering, with a focus on building production data pipelines and data products.
  • Advanced SQL and Python; hands-on Databricks / Unity Catalog fluency (notebooks, Delta Lake, catalog/schema/grants, Workflows).
  • Qualified applicants must be authorized to work in the United States on a full-time basis. Lilly will not provide support for or sponsor work authorization or visas for this role, including but not limited to F-1 CPT, F-1 OPT, F-1 STEM OPT, J-1, H-1B, TN, O-1, E-3, H-1B1, or L-1.

Additional Skills / Preferences

  • Experience defining and executing data ingestion pipelines at enterprise scale.
  • Working knowledge of data governance, classification, and access control (RBAC/ABAC, row- and column-level security).
  • Proficiency with Git-based CI/CD workflows (Git Actions or equivalent) for versioned data and pipeline artifacts.
  • Excellent problem-solving skills; ability to translate architectural designs into working, tested pipelines.
  • Good communication and collaboration skills — able to work effectively with architects, product owners, and multi-functional partners.
  • Experience in regulated or healthcare data environments; familiarity with HIPAA, PHI handling, and covered-entity constructs.
  • Data contract frameworks (ODCS or comparable) and policy-as-code (OPA/Rego) awareness.
  • Tokenization / de-identification pattern implementation (Datavant or comparable).
  • PySpark and transformation-as-code (DLT); test frameworks (pytest, DQX / Great Expectations).
  • Infrastructure-as-code (Terraform) for provisioning Unity Catalog objects and grants.
  • MDM / identity resolution implementation experience (Reltio or comparable).
  • Experience with Agile/Scrum methodologies and Jira.
  • Experience building or integrating AI Skills and agents against governed data platforms.

Similar jobs

Senior Data Engineer

Apex SystemsLincoln, NE· 1 mo ago
Information Technologyapply on apexsystems.com

Senior Data Engineer

OptumEden Prairie, MN· 2 mo ago
Information Technology$92k–$164k/yrapply on careers.unitedhealthgroup.com

Senior Data Engineer

QuantiphiBoston, MA· 2 mo ago
Information Technologyapply on quantiphi.wd1.myworkdayjobs.com

Senior Data Engineer

Gen II Fund ServicesBoca Raton, FL· 2 mo ago
Information Technology$140k/yrapply on recruiting.ultipro.com

Senior Data Engineer

Boost MobileLittleton, CO· 1 mo ago
Information Technology$148k–$158k/yrapply on jobs.echostar.com

Senior Data Engineer

AllianceBernsteinNashville, TN· 1 mo ago
Information Technologyapply on abglobal.wd1.myworkdayjobs.com