Jobs · Information Technology · Indiana

Data Engineer - Lilly Medicine Foundry

BioSpace · Indianapolis, IN · 4 days ago
Information Technology$126k–$224k/yrFull-time

What You'll Be Doing

As a Data Engineer based at the Lilly Medicine Foundry, in Lebanon, IN, you will design and build data pipelines for the site. Working closely with the Data Architect and Data Scientists, you'll engage directly with business and operational collaborators to understand how data flows across the Foundry — from lab systems, process engineering, and manufacturing execution to quality and process analytics — and translate that understanding into reliable, scalable data products.

Your core focus will be integrating IT and OT source systems with cloud data Lakehouse architecture spanning AWS and Azure, enabling the advanced analytics and AI/ML capabilities that will define how Foundry operates. This includes the full data value chain: capture, ingestion, integration, contextualization, harmonization, and delivery as reusable data domains and data-as-a-product.

You will work across enterprise and edge systems, navigate a diverse technology landscape, and ensure that everything you build meets the data integrity and compliance standards required in a GxP-regulated manufacturing environment.

How You'll Succeed

  • Build and deliver high-quality data solutions:

    • Design and develop robust data pipelines for ingestion, transformation, and integration across cloud platforms (AWS and Azure), applying modern patterns such as medallion architecture, streaming ingestion, and API-based extraction.
    • Write clean, maintainable code, optimize for performance, and think carefully about how data flows from source systems into trustworthy, analytics-ready products.
  • Translate business needs into technical outcomes:

    • Engage proactively with stakeholders across multiple modalities and business teams including Operations, Quality, Lab, and Engineering to understand their data needs.
    • Identify gaps, translate operational realities into sound technical designs, and explain your architectural decisions back to a non-technical audience.
  • Navigate a diverse enterprise landscape:

    • Invest time in understanding upstream and downstream dependencies, seek opportunities to reuse existing services and integrations, and build with the broader data architecture in mind.
  • Ensure compliance:

    • Understand the data integrity requirements of a GxP manufacturing environment and apply Quality guidelines and Good Manufacturing Practices from the start.
    • Participate in design reviews, maintain traceability, and contribute to governance frameworks that keep the Foundry consistently compliant.
  • Continuous learning:

    • Track emerging tools and technologies across the AWS and Azure data ecosystems, bring informed perspectives on what's worth adopting, and actively contribute to the team's collective capability.

Basic Qualifications

  • Bachelor’s degree in Computer Science, Data Science, Engineering or related field
  • At least 3 years of experience in several of the following disciplines: statistical methods, data modeling, ETL/ELT, ontology development, semantic graph construction and linked data, relational schema design.
  • Experience with cloud platforms (e.g., AWS, Azure).
  • Experience with AI/ML/LLM Concepts and tools and building agentic AI solution sets.

Additional Skills / Preferences

  • At least 1 year of experience in a pharmaceutical GxP or Scientific environment.
  • 1-3 years of experience designing large scale data models for functional, operational, and analytical environments (Conceptual, Logical, Physical & Dimensional).
  • Demonstrated SQL and data modeling proficiency.
  • Experience with data modeling tools such as, ER*Studio and Erwin or TOAD.
  • Experience with data integration such as data streaming, Industrial IOT, using MQTT, AQMP, Kafka and related protocols.
  • Understanding of modern data architecture, data lakehouse, data warehousing and/or big data concepts.
  • Experience with security models and development on large data sets.
  • Experience with multiple database solutions (e.g. Postgres, Redshift, Aurora, Athena, Graph DB like Neptune, No SQL like DynamoDB, MongoDB) and formal database designs (3NF, Dimensional Models).
  • Experience with Agile Development, CI/CD, Github, Automation platforms.
  • Demonstrated ability to analyze large, complex data domains and craft practical solutions for subsequent data exploitation via analytics.
  • Knowledgeable in data functions such as Data Governance, Master Data Management, Business Intelligence.
  • Prior work experience working in pharma or other GMP setting.
  • Solid knowledge of Computer System Validation process.
  • Demonstrated ability to analyze, anticipate, and resolve complex issues through sound problem-solving skills.
  • Demonstrated learning agility and curiosity.
  • Desire and ability to communicate using a variety of methods in diverse forums.

Similar jobs