Senior Software Engineer — Lakehouse Systems
About the role
Granica builds AI infrastructure for enterprises operating massive data environments. This role involves building foundational lakehouse systems for AI.
Responsibilities
- Build metadata and transaction systems for large-scale tabular datasets
- Design systems that support time travel, schema evolution, partition evolution, snapshot isolation, and atomic consistency
- Develop table-maintenance infrastructure for lakehouse formats such as Apache Iceberg, Delta Lake, and Apache Hudi
- Build systems for manifests, snapshots, transaction logs, metadata pruning, snapshot expiration, table garbage collection, and catalog consistency
- Optimize file layout, clustering, compaction, file sizing, data skipping, indexing, and read-path performance
- Improve performance and cost efficiency across object-store-backed lakehouse environments such as S3, GCS, and ADLS
- Work with columnar formats such as Parquet and ORC, including encoding, compression, layout, pruning, and read-path optimization
- Build systems that make lakehouse tables faster, cheaper, and more reliable across engines and platforms such as Spark, Flink, Trino, Presto, Databricks, and Snowflake-adjacent environments
- Debug performance bottlenecks across storage, metadata, table maintenance, query execution, network, and compute layers
- Develop workload-aware table optimization systems that learn from access patterns and reorganize data automatically
- Implement algorithms in compression, representation, layout optimization, and data efficiency
- Contribute to open-source or publish research when appropriate
Requirements
Strong engineering depth in distributed systems, storage systems, databases, or data infrastructure
Production experience with modern data lake or lakehouse technologies such as Iceberg, Delta Lake, Hudi, Spark, Trino, Presto, Flink, Hive Metastore, Unity Catalog, or similar systems
Hands-on experience with columnar formats such as Parquet or ORC
Understanding of metadata-driven architectures, table formats, transaction semantics, query planning, and physical data layout
Experience with table maintenance, compaction, clustering, file sizing, metadata pruning, snapshot expiration, or garbage collection
Familiarity with cloud object storage systems such as S3, GCS, or ADLS and the performance tradeoffs of building lakehouse systems on top of them
Strong programming skills in Java, Scala, Go, Rust, C++, or similar systems-oriented languages
Curiosity about compression, entropy, information theory, and how data representation affects AI efficiency
A pragmatic builder’s mindset: rigorous, hands-on, and comfortable owning complex systems end to end
Qualifications
Experience contributing to Apache Iceberg, Delta Lake, Apache Hudi, Spark, Flink, Trino, Presto, Velox, DuckDB, Polars, Parquet, ORC, or related systems
Experience with manifests, snapshots, metadata catalogs, schema evolution, partition evolution, delete handling, transaction logs, or table garbage collection
Experience solving the small-file problem, optimizing object-store access patterns, or improving table health at scale
Skills
Storage engines, query engines, indexing, caching, encoding, compression, or adaptive query optimization
Benefits
Competitive salary, meaningful equity, and performance bonus for top performers
401(k) with company match, comprehensive health coverage, and unlimited PTO
Daily catered meals in our Mountain View office
Support for research, publication, and conference participation