Senior Software Engineer — Lakehouse Systems
Granica · San Francisco Bay Area · 2 days ago
On-siteEngineering$160k–$240k/yrFull-time
About the role
Granica is hiring a Senior Software Engineer to build foundational lakehouse systems for AI. You will work on the core infrastructure behind Crunch, Granica’s continuous optimization product for enterprise lakehouse data.
Responsibilities
- Build metadata and transaction systems for large-scale tabular datasets
- Design systems that support time travel, schema evolution, partition evolution, snapshot isolation, and atomic consistency
- Develop table-maintenance infrastructure for lakehouse formats such as Apache Iceberg, Delta Lake, and Apache Hudi
- Create systems for manifests, snapshots, transaction logs, metadata pruning, snapshot expiration, table garbage collection, and catalog consistency
- Optimize file layout, clustering, compaction, file sizing, data skipping, indexing, and read-path performance
- Improve performance and cost efficiency across object-store-backed lakehouse environments such as S3, GCS, and ADLS
- Work with columnar formats such as Parquet and ORC, including encoding, compression, layout, pruning, and read-path optimization
- Build systems that make lakehouse tables faster, cheaper, and more reliable across engines and platforms such as Spark, Flink, Trino, Presto, Databricks, and Snowflake-adjacent environments
- Debug performance bottlenecks across storage, metadata, table maintenance, query execution, network, and compute layers
- Develop workload-aware table optimization systems that learn from access patterns and reorganize data automatically
- Implement algorithms in compression, representation, layout optimization, and data efficiency
- Contribute to open-source or publish research when appropriate
Requirements
- Strong engineering depth in distributed systems, storage systems, databases, or data infrastructure
- Production experience with modern data lake or lakehouse technologies such as Iceberg, Delta Lake, Hudi, Spark, Trino, Presto, Flink, Hive Metastore, Unity Catalog, or similar systems
- Hands-on experience with columnar formats such as Parquet or ORC
- Understanding of metadata-driven architectures, table formats, transaction semantics, query planning, and physical data layout
- Experience with table maintenance, compaction, clustering, file sizing, metadata pruning, snapshot expiration, or garbage collection
- Familiarity with cloud object storage systems such as S3, GCS, or ADLS and the performance tradeoffs of building lakehouse systems on top of them
- Strong programming skills in Java, Scala, Go, Rust, C++, or similar systems-oriented languages
- Curiosity about compression, entropy, information theory, and how data representation affects AI efficiency
- A pragmatic builder’s mindset: rigorous, hands-on, and comfortable owning complex systems end to end
Qualifications
- Experience contributing to Apache Iceberg, Delta Lake, Apache Hudi, Spark, Flink, Trino, Presto, Velox, DuckDB, Polars, Parquet, ORC, or related systems
- Experience with manifests, snapshots, metadata catalogs, schema evolution, partition evolution, delete handling, transaction logs, or table garbage collection
- Experience solving the small-file problem, optimizing object-store access patterns, or improving table health at scale
- Background in storage engines, query engines, indexing, caching, encoding, compression, or adaptive query optimization
- Research or open-source contributions in distributed systems, databases, storage, compression, indexing, or data processing
- Interest in how physical data representation affects model training, inference, retrieval, and reasoning efficiency
Skills
- Strong engineering depth in distributed systems, storage systems, databases, or data infrastructure
- Production experience with modern data lake or lakehouse technologies such as Iceberg, Delta Lake, Hudi, Spark, Trino, Presto, Flink, Hive Metastore, Unity Catalog, or similar systems
- Hands-on experience with columnar formats such as Parquet or ORC
- Understanding of metadata-driven architectures, table formats, transaction semantics, query planning, and physical data layout
- Experience with table maintenance, compaction, clustering, file sizing, metadata pruning, snapshot expiration, or garbage collection
- Familiarity with cloud object storage systems such as S3, GCS, or ADLS and the performance tradeoffs of building lakehouse systems on top of them
- Strong programming skills in Java, Scala, Go, Rust, C++, or similar systems-oriented languages
- Curiosity about compression, entropy, information theory, and how data representation affects AI efficiency
- A pragmatic builder’s mindset: rigorous, hands-on, and comfortable owning complex systems end to end
Benefits
- Competitive salary, meaningful equity, and performance bonus for top performers
- 401(k) with company match, comprehensive health coverage, and unlimited PTO
- Daily catered meals in our Mountain View office
- Support for research, publication, and conference participation
Pay
Compensation Range: $160K - $240K
Schedule
On-site