Jobs · Information Technology · New York

Staff Data Engineer

Butterfly Network, Inc. · New York, NY · 2 wk ago
Information TechnologyFull-time

Butterfly Network, Inc. (NYSE: BFLY) is driving a digital revolution in ultrasound imaging and sensing with its proprietary Ultrasound-on-Chip™ semiconductor technology and software solutions. Butterfly first proved its technology in the point-of-care ultrasound market – commercializing the world’s first single-probe, whole-body portable ultrasound device, now on its best-selling, third-generation: Butterfly iQ3™. The Company combines its advanced hardware with cloud software and AI, an enterprise workflow solution (Compass AI™), and other offerings to drive adoption of affordable, accessible ultrasound. Butterfly also enables third-party development of imaging AI apps through Butterfly Garden™, its software development kit and AI partnership initiative. In addition, Butterfly Embedded™ is the Company’s Ultrasound-on-Chip™ licensing and co-development program designed to enable novel ultrasound applications across non-competitive healthcare markets and beyond.

About the role

Butterfly Network is seeking a Staff Data Engineer to build and own the data layers that power every dashboard, data product, and AI agent we build. This is not a ticket-executor role: we need someone who designs the solution first, then builds it. You’ll work on a small, high-ownership team directly under the Director of Data and AI Platforms, with broad scope and real accountability for the quality and reliability of data that the rest of the company depends on. The team is migrating from a GCP-based platform supporting BI to an AWS/Databricks platform supporting internal and external-facing agents while continuing to support traditional BI.

You’ll pick up active migration workstreams, take ownership of ingestion pipelines and dbt models, and help establish the engineering patterns that will carry the platform forward. The near-term work includes GCP → AWS cutover, new source onboarding, development of models in medallion architecture, and reverse ETL flows back into operational systems.

Responsibilities

  • Own ingestion pipelines end-to-end. Design and implement reliable, observable pipelines from operational sources into the Bronze layer using CDC patterns, Autoloader/Lakeflow, and batch ingestion.
  • Own the transformation layer. Build and maintain dbt models in a medallion architecture, set up testing and alerting, document data assets in Unity Catalog, and ensure sensitive data is tagged and access-controlled.
  • Drive GCP to AWS/Databricks migration. Take ownership of active pipeline cutovers from Airflow, Cloud Run, Cloud Functions, and BigQuery to Databricks on AWS. Validate parity, coordinate with stakeholders, and decommission legacy components without disrupting business-critical reporting.
  • Build and maintain reverse ETL and operational integrations. Sync curated data back into Salesforce, NetSuite, and other operational systems via Mulesoft.
  • Architect before building. Write ADRs for non-trivial design decisions. Define patterns the team follows for ingestion, transformation, and data quality. Make deliberate build-vs-managed tradeoffs and document the reasoning.
  • Translate business needs into data requirements. Partner with Sales, Finance, Marketing, Clinical, and Product teams to understand their data workflows and source systems. Convert that context into well-scoped pipeline and model work with clear acceptance criteria.

Requirements

  • 5+ years of data engineering experience, with clear evidence of seniority: owned complex domains, designed solutions from scratch, and made architectural decisions — not just executed tickets.
  • Strong hands-on dbt experience: models, tests, sources, macros, documentation, CI integration, and refactoring existing work.
  • Demonstrated experience with both streaming and batch ingestion patterns, including CDC pipelines, event-driven architectures, and scheduled bulk loads from operational sources.
  • Experience with Databricks (Delta Lake, Delta Live Tables) or a comparable lakehouse platform.
  • Solid Python for data engineering: PySpark, pipeline development, utilities, and custom tooling.
  • Hands-on experience integrating with operational data sources: CRM (Salesforce), ERP (NetSuite), payments (Stripe), or similar.
  • Ability to design before building: write the doc, define the interface, identify the failure modes, then implement.
  • Experience with data quality, testing, and pipeline observability: dbt tests, Great Expectations, alerting, SLA tracking.

Qualifications

Ideally, you also have these skills/experiences/attributes (but it’s ok if you don’t!):

  • Experience with Unity Catalog or an equivalent governance layer for column masking, row-level security, lineage, and sensitivity tags.
  • Familiarity with reverse ETL or integration platforms (Mulesoft, Boomi, or equivalent).
  • IaC experience at the data workload level (Terraform, Spacelift, or equivalents).
  • Experience in healthcare, life sciences, or another regulated environment (HIPAA, PII/PHI classification and handling).
  • Exposure to AI/ML platform patterns: vector search, RAG pipelines, or model serving data flows.

Values

  • Patient-Centric Innovators: Our mission is THE mission.
  • Empowered to Impact: Every voice matters.
  • One Team, One Goal: Unity fuels progress.
  • Growth Champions: We embrace challenges.
  • Action-Oriented Achievers: We follow through, every time.

Location

Butterfly offers a hybrid work model for most positions, with team members spending two or more days a week in the office. While flexibility is key, we value in-person connections that spark creativity and teamwork. Our offices are designed for collaboration, with comfortable workspaces, stocked kitchens, and opportunities to connect with peers. This is a hybrid position that can be based in San Francisco, CA or New York City, NY.

Benefits

  • Comprehensive health insurance, including dental and vision coverage.
  • Annual contributions to Health Savings Account (HSA) for all enrolled employees.
  • Comprehensive Employee Assistance Program to support emotional health and day-to-day needs.
  • 401k plan with company match.
  • Employee Stock Purchase Plan (ESPP) for eligible employees.
  • Unlimited Paid Time Off + 10 Holiday Days a Year.
  • Parental Leave with additional support for primary caregivers transitioning back to work.
  • Competitive salaried compensation.
  • Equity to make every employee a stakeholder.
  • The opportunity to build a revolutionary healthcare product and save millions of lives.

Pay

Our estimated salary for this role is around $185,000 - $195,000 + bonus + equity + benefits. Actual pay is determined by multiple factors such as skills, qualifications, experience, and market demand.

For this role, we are only considering candidates who are legally authorized to work in the United States and who do not now or in the future require sponsorship for employment visa status.

Similar jobs

Staff Data Engineer

Necessary VenturesBerkeley, CA· 1 mo ago
Information Technologyapply on copperhome.com

Staff Data Engineer

HCA HealthcareNashville, TN· 2 mo ago
Healthcareapply on careers.hcahealthcare.com

Staff Data Engineer

Awetomaton LtdSt Louis, MO· 1 mo ago
Information Technology$25k/yrapply on grnh.se

Staff Data Engineer

Self Financial, Inc.Austin, TX· 2 mo ago
Information Technology$160k–$195k/yrapply on grnh.se

Staff Data Engineer

ImprintBuffalo-Niagara Falls Area· 2 mo ago
Information Technologyapply on jobs.ashbyhq.com

Staff Data Engineer

XometryNorth Bethesda, MD· 2 mo ago
Information Technology$180k–$200k/yrapply on job-boards.greenhouse.io