Principal Data Platform Architect With Strong Iceberg/Trino
Novia Infotech · United States · 1 mo ago
RemoteRemoteDesignContract
About the role
Own the lakehouse reference architecture including Iceberg table design, Trino cluster topology, catalog service, Spark transform compute, and object-storage layout. Design on-premise replacements for cloud-managed warehouse capabilities with no direct equivalent, including change-data-capture streams, scheduled tasks, and write-back paths into operational stores. Lead the SQL dialect strategy for porting existing warehouse workloads to Trino and Spark SQL.
Responsibilities
- Own the lakehouse reference architecture: Iceberg table design, Trino cluster topology, catalog service, Spark transform compute, and object-storage layout
- Design on-premise replacements for cloud-managed warehouse capabilities that have no direct equivalent: change-data-capture streams, scheduled tasks, and write-back paths into operational stores
- Run proof-of-concept validation of the catalog and query engine at expected data volumes, and define evidence-based triggers for placement decisions (VM-based versus Kubernetes-native operators)
- Set platform-wide standards for table layout, partitioning, file sizing, and Iceberg maintenance: compaction, snapshot expiry, and orphan-file cleanup
- Lead the SQL dialect strategy for porting existing warehouse workloads to Trino and Spark SQL
- Mentor senior engineers across data workstreams, review designs, and raise the bar on engineering quality
- Partner with platform engineering on storage sizing, resource isolation, and capacity planning for the lakehouse footprint
Requirements
- B.E., B.Tech., M.Sc. degree in Computer Science or a related technical field
- 12+ years of industry experience building and operating large-scale data platforms or distributed systems
- Deep, hands-on expertise with distributed SQL engines: Trino/Presto or Spark SQL internals, query planning, and performance engineering
- Production experience with Apache Iceberg (or Delta Lake/Hudi with willingness to go deep on Iceberg): table spec, merge-on-read versus copy-on-write, and table maintenance at scale
- Working knowledge of Iceberg catalog services (REST catalogs such as Polaris or Nessie, or Hive Metastore) and S3-compatible object storage
- Strong understanding of cloud warehouse internals (Snowflake, BigQuery, or Redshift) sufficient to design functional equivalents on open-source infrastructure
- Professional software development experience with Java and/or Python
Qualifications
- Experience delivering data platforms in on-premise, regulated, or air-gapped environments is a strong plus
- Healthcare data experience is a plus