Jobs · Information Technology · California

Staff Data Engineer

Butterfly Network, Inc. · California, United States · Today
Information TechnologyFull-time

About the Role

Butterfly Network is seeking a Staff Data Engineer to build and own the data layers that power every dashboard, data product, and AI agent we build. This is not a ticket-executor role: we need someone who designs the solution first, then builds it. You'll work on a small, high-ownership team directly under the Director of Data and AI Platforms, with broad scope and real accountability for the quality and reliability of data that the rest of the company depends on.

The team is in the middle of a full migration from a GCP-based platform that supported BI to an AWS/Databricks platform supporting internal and external facing agents while continuing to support traditional BI. You'll pick up active migration workstreams, take ownership of ingestion pipelines and dbt models, and help establish the engineering patterns that will carry the platform forward. The near-term work is concrete: GCP → AWS cutover, new source onboarding, development of models in medallion architecture, and reverse ETL flows back into operational systems.

Core Responsibilities

  • Own ingestion pipelines end-to-end. Design and implement reliable, observable pipelines from operational sources into the Bronze layer using CDC patterns, Autoloader/Lakeflow, and batch ingestion.
  • Own the transformation layer. Build and maintain dbt models in a medallion architecture, set up testing and alerting, document data assets in Unity Catalog, and ensure sensitive data is tagged and access-controlled.
  • Drive GCP to AWS/Databricks migration. Take ownership of active pipeline cutovers from Airflow, Cloud Run, Cloud Functions, and BigQuery to Databricks on AWS. Validate parity, coordinate with stakeholders, and decommission legacy components without disrupting business-critical reporting.
  • Build and maintain reverse ETL and operational integrations. Sync curated data back into Salesforce, NetSuite, and other operational systems via Mulesoft.
  • Architect before building. Write ADRs for non-trivial design decisions. Define patterns the team follows for ingestion, transformation, and data quality. Make deliberate build-vs-managed tradeoffs and document the reasoning.
  • Translate business needs into data requirements. Partner with Sales, Finance, Marketing, Clinical, and Product teams to understand their data workflows and source systems. Convert that context into well-scoped pipeline and model work with clear acceptance criteria.

Qualifications

Baseline skills/experiences/attributes:

  • 5+ years of data engineering experience, with clear evidence of seniority: owned complex domains, designed solutions from scratch, and made architectural decisions — not just executed tickets.
  • Strong hands-on dbt experience: models, tests, sources, macros, documentation, CI integration, and refactoring existing work.
  • Demonstrated experience with both streaming and batch ingestion patterns, including CDC pipelines, event-driven architectures, and scheduled bulk loads from operational sources.
  • Experience with Databricks (Delta Lake, Delta Live Tables) or a comparable lakehouse platform.
  • Solid Python for data engineering: PySpark, pipeline development, utilities, and custom tooling.
  • Hands-on experience integrating with operational data sources: CRM (Salesforce), ERP (NetSuite), payments (Stripe), or similar.
  • Ability to design before building: write the doc, define the interface, identify the failure modes, then implement.
  • Experience with data quality, testing, and pipeline observability: dbt tests, Great Expectations, alerting, SLA tracking.

Ideally, you also have these skills/experiences/attributes (but it's ok if you don't!):

  • Experience with Unity Catalog or an equivalent governance layer for column masking, row-level security, lineage, and sensitivity tags.
  • Familiarity with reverse ETL or integration platforms (Mulesoft, Boomi, or equivalent).
  • IaC experience at the data workload level (Terraform, Spacelift, or equivalents).
  • Experience in healthcare, life sciences, or another regulated environment (HIPAA, PII/PHI classification and handling).
  • Exposure to AI/ML platform patterns: vector search, RAG pipelines, or model serving data flows.

Location

Butterfly offers a hybrid work model for most positions, with team members spending two or more days a week in the office. While flexibility is key, we value in-person connections that spark creativity and teamwork. Our offices are designed for collaboration, with comfortable workspaces, stocked kitchens, and opportunities to connect with peers. This is a hybrid position that can be based in San Francisco, CA or New York City, NY.

Benefits and Perks

  • Comprehensive health insurance, encompassing dental and vision coverage, is provided to all our employees. As a health-tech company, we prioritize the well-being of our teams. We also contribute to Health Savings Account (HSA) accounts for all enrolled employees on an annual basis.
  • Comprehensive Employee Assistance Program - we provide access to tools and resources to support your emotional health and day-to-day needs.
  • 401k plan and match - we facilitate your retirement goals.
  • Eligible employees will have the opportunity to participate in Employee Stock Purchase Plan (ESPP).
  • Unlimited Paid Time Off + 10 Holiday Days a Year - recharge and come back ready to make an impact.
  • Parental Leave - we aim to provide our employees with time to bond with their growing family, along with additional support for primary caregivers to help transition back to work.
  • Competitive salaried compensation - we value our employees and show it.
  • Equity - we want every employee to be a stakeholder.
  • The opportunity to build a revolutionary healthcare product and save millions of lives!

Pay

Our estimated salary for this role is around $185,000 - $195,000 + bonus + equity + benefits. Actual pay is determined by multiple factors such as skills, qualifications, experience and market demand.

Similar jobs

Staff Data Engineer

TwentyNew York, NY· 1 mo ago
Information Technology$138/hrapply on jobs.ashbyhq.com

Staff Data Engineer

ModMedBoca Raton, FL· 2 wk ago
Information Technologyapply on modmed.wd501.myworkdayjobs.com

Staff Data Engineer

Flywheel Energy, LLCOklahoma City, OK· 1 mo ago
Information Technologyapply on paycomonline.net

Staff Data Engineer

LyricUnited States· 1 wk ago
RemoteInformation Technology$104k–$157k/yrapply on tbc.wd12.myworkdayjobs.com

Staff Data Engineer

IntuitSan Diego, CA· 1 wk ago
Information Technology$192k–$259k/yrapply on dsp.prng.co

Staff Data Engineer

Checkr, Inc.Denver, CO· 1 wk ago
Information Technology$196k–$230k/yrapply on job-boards.greenhouse.io