Senior Data Infrastructure Engineer
Aircall · San Francisco, CA · 2 days ago
HybridEngineering$150k/yrFull-time
About the role
Aircall's Data team is mid-migration: we are moving off a single Redshift cluster onto an Apache Iceberg lakehouse on S3, with Flink CDC into Kafka for ingestion and dbt-on-Spark via Apache Kyuubi on EKS for transformation. It's a real greenfield platform build — already scoped and underway — on top of a stack that carries ten years of startup-growth history, and all the quirks that come with it. We're building this role to give platform work the runway it deserves.
Responsibilities
- Build and operate the lakehouse: Apache Iceberg on S3, table design and maintenance, partitioning and compaction, and the migration of remaining Redshift workloads onto it
- Own ingestion end to end — Flink CDC → Kafka (MSK) → Iceberg, plus Rudderstack, Fivetran and DMS sources — and hold the freshness and reliability SLAs on it
- Run and evolve the compute and orchestration layer: Apache Kyuubi on EKS for dbt-spark, Airflow (completing its ECS → EKS migration), autoscaling, spot strategy and cost efficiency
- Build the tooling, libraries and templates that let analytics engineers and data scientists own their own pipelines without filing a ticket — self-service is the deliverable, not a side effect
- Close our environment gaps: a real staging environment, CI that tests against staging rather than production, automated schema-change detection, gated promotion and canary deploys for critical models
- Own governance and access at the platform level: Lake Formation row/column RBAC, StrongDM zero-trust access, SSO, audit logging, and PII handling
- Own observability: Monte Carlo, lineage, alerting and the SLAs we publish — and drive incidents to root cause and to a durable fix
- Champion infrastructure as code and automation (Terraform, GitLab CI, GitOps) across everything the team runs
Qualifications
- 4+ years (Senior: 6+) in data engineering, data platform or infrastructure engineering
- Strong Python and SQL, with demonstrated experience building frameworks and tooling others depend on, not only pipelines
- Production experience with an orchestration framework (Airflow, Dagster, Prefect) at meaningful scale — including the operational side, not just DAG authoring
- Deep AWS experience (S3, EKS/ECS, IAM, Glue/Athena or equivalent)
- Comfortable building and debugging CI/CD, infrastructure as code (Terraform) and GitOps workflows; familiar with Kubernetes and Docker
- Track record owning reliability: SLAs, monitoring, alerting, on-call, and post-incident hardening
- Daily, hands-on use of AI coding tools (Claude Code, Cursor, or equivalent) as a core part of how you build and operate infrastructure
- Great cross-functional communication — you'll shape data contracts with backend engineering and align expectations with analytics consumers