Senior Manager, Data Engineering
About the Role
You will own how data gets into our platform and how it gets served back out — ingestion, the lakehouse, and the query layer underneath everything analytics and product depend on. The problem is specific: data arrives from CDC streams, transactional databases, event topics, partner APIs, and files, and today each source carries its own pipeline, its own failure modes, and its own on-call story. Your mandate is to collapse that into one ingestion framework and one open lakehouse — reliable enough to publish SLOs against, fast enough to serve interactive query, and cheap enough to defend line by line.
This is a hands-on role. You will set technical direction, hire, and grow the team — and you will also be in design reviews, in code review, and in the pipeline when a stateful stream will not recover from its checkpoint. Expect roughly half your time in technical work. You own the full path: how data lands, the table format and its lifecycle, the transformation layer, the engines that serve it, and the SLOs on top of all of it. When a dataset is late or wrong, it is your team's phone that rings — and you are expected to have already built the thing that catches it first.
Responsibilities
- Lead, hire, and grow a team of 10+ data engineers — set the technical bar through design and code review, not through status meetings.
- Own the architecture and delivery of a unified ingestion framework: one configuration-driven path for batch, CDC, and streaming sources, with schema evolution, replay and backfill, idempotency, dead-letter handling, and data contracts built into the framework rather than reimplemented per pipeline.
- Own production Spark Structured Streaming pipelines — watermarking, stateful joins and aggregations, checkpoint and restart discipline, exactly-once sinks, lag and backpressure management.
- Own the Apache Iceberg lakehouse: partition and sort strategy, file sizing and compaction, snapshot and orphan-file lifecycle, schema and partition evolution, and multi-engine interoperability.
- Set the dbt modeling standard — layering conventions, tests, contracts, CI enforcement, and lineage that stakeholders trust.
- Own Trino catalog design, workload isolation, and query performance for interactive and federated access.
- Define and meet freshness, completeness, and latency SLOs. Run a 24x7 on-call rotation with a short mean time to restore.
- Own cost: a defensible cost-per-pipeline and cost-per-dataset number, and the levers to move it.
- Partner with product, analytics, and architecture to sequence the roadmap, and bring rigor to decisions — collect the data, seek dissent, and run pilots rather than arguing from opinion.
Requirements
- Bachelor's degree in Computer Science, Engineering, or equivalent practical experience.
- 8+ years in data engineering, including 3+ years leading engineers as a manager or tech lead — and you are still hands-on in code and design.
- Production experience with Spark Structured Streaming at scale: state store growth, checkpoint recovery, watermark tuning, and late or out-of-order data.
- Deep Apache Spark and PySpark performance work — diagnosing and fixing skew, shuffle pressure, small-file problems, and executor memory failures on real workloads.
- Experience building or substantially owning a reusable ingestion framework serving multiple source types — not a collection of individual pipelines.
- Production experience with an open table format (Apache Iceberg preferred) including schema and partition evolution, compaction strategy, and migration from an existing format.
- Experience with dbt as a team-wide modeling standard, including testing and CI.
- Experience with Trino or Presto operations and query optimization.
- Experience running reliable, high-scale platform systems — 24x7 on-call, availability targets, and fast restoration of service.
Preferred Qualifications
- Apache Flink, Kafka or MSK internals, or high-throughput stream-join design.
- CDC tooling in production (Debezium, DMS, GoldenGate, Qlik) and integrating legacy or mainframe sources into modern pipelines.
- Data contracts, catalog, and lineage tooling (DataHub, OpenMetadata, Glue, Unity).
- Iceberg REST catalog implementations and multi-engine interoperability.
- AWS, Kubernetes, and Terraform fluency — you can debug below the framework layer.
- Multi-tenant B2B data platforms with per-tenant cost attribution and isolation.
- Open-source contribution to the projects in this stack.
What Success Looks Like
- A stable, well-led SRE organization with clear ownership, career paths, and low regrettable attrition among your managers and their teams.
- Consistent, Datadog-driven observability and SLOs in place across the organization, with measurable reduction in Sev1/Sev2 incidents and mean time to detect/resolve.
- Modern, standardized infrastructure practices — GitOps delivery via Argo, IaC via Terraform/OpenTofu, and reliable CI/CD via Jenkins — adopted consistently across teams and clouds.
- A mature, blameless incident-management culture with strong postmortem follow-through.
- Strong cross-functional trust with engineering, product, and security/compliance stakeholders.
Benefits
- Competitive salary and market-competitive benefits.
- Quarterly perks program.
- Ample time off and hybrid working arrangements where appropriate.
- Access to continuous learning, professional certifications, and leadership development opportunities.
- Supportive global culture fostering diversity, inclusion, and values authenticity, trust, curiosity, and diversity of thought.