Jobs · Information Technology · New York

Staff Data Engineer - Data Engineering

CoreWeave · New York, NY · Yesterday
Information Technology$207k–$275k/yrFull-time

About the role

The Data Engineering Team builds and operates the foundational data infrastructure powering analytics, AI, and operational decision-making across CoreWeave. We design resilient data pipelines, scalable lakehouse systems, and high-quality datasets that enable teams across Finance, HR, Operations, and Engineering to move faster and make smarter decisions.

Responsibilities

  • Define and drive the architecture of CoreWeave's enterprise data ecosystem.
  • Establish the modeling, semantic, governance, and platform standards that enable teams to build trusted and reusable data products at scale.
  • Lead cross-domain architecture, solve the organization's most complex data design challenges, and partner with senior engineers and stakeholders to translate business requirements into durable technical systems.
  • Combine deep hands-on engineering expertise with company-wide technical leadership across our modern lakehouse platform.
  • Lead architecture reviews for high-impact initiatives and ensure alignment with the broader platform strategy.
  • Evaluate and standardize core data technologies across processing, orchestration, cataloging, modeling, and metadata management.
  • Mentor senior engineers and raise the quality of system design and technical decision-making across the organization.

Requirements

10+ years of experience in data engineering, software engineering, distributed systems, or data architecture roles. Demonstrated experience setting technical direction for data systems spanning multiple teams, business domains, or platforms. Deep expertise in enterprise analytical modeling, including dimensional, Data Vault, semantic modeling, and governed metrics. Experience designing and operating large-scale lakehouse or streamhouse architectures using technologies such as Apache Iceberg, Delta Lake, Apache Hudi, Apache Paimon, or Apache Fluss. Advanced knowledge of distributed OLAP, query, processing, and ingestion systems such as StarRocks, ClickHouse, Trino, Spark, Flink, or Kafka. Expert-level SQL and strong programming expertise in Python, Scala, Java, or Rust, with experience building production-grade data systems. Demonstrated experience implementing governance capabilities such as lineage, profiling, metadata management, data quality controls, access policies, or auditability. Demonstrated ability to independently turn ambiguous problem statements into clear technical direction and execution plans. Track record of leading cross-cutting architecture and establishing standards adopted across multiple engineering teams. Preferred Experience: Designing data systems that support compliance programs such as SOX, GDPR, MNPI, and PII controls. Experience operating event-driven architectures using Kafka, AutoMQ, Pulsar, or similar streaming platforms. Experience running data-intensive workloads on Kubernetes. Experience in high-growth cloud, infrastructure, or distributed-systems organizations.

Similar jobs