Jobs · Engineering · California

Senior Software Engineer — Lakehouse Systems

Granica · San Francisco Bay Area · 2 days ago
On-siteEngineering$160k–$240k/yrFull-time

About the role

Granica is hiring a Senior Software Engineer to build foundational lakehouse systems for AI. You will work on the core infrastructure behind Crunch, Granica’s continuous optimization product for enterprise lakehouse data.

Responsibilities

  • Build metadata and transaction systems for large-scale tabular datasets
  • Design systems that support time travel, schema evolution, partition evolution, snapshot isolation, and atomic consistency
  • Develop table-maintenance infrastructure for lakehouse formats such as Apache Iceberg, Delta Lake, and Apache Hudi
  • Create systems for manifests, snapshots, transaction logs, metadata pruning, snapshot expiration, table garbage collection, and catalog consistency
  • Optimize file layout, clustering, compaction, file sizing, data skipping, indexing, and read-path performance
  • Improve performance and cost efficiency across object-store-backed lakehouse environments such as S3, GCS, and ADLS
  • Work with columnar formats such as Parquet and ORC, including encoding, compression, layout, pruning, and read-path optimization
  • Build systems that make lakehouse tables faster, cheaper, and more reliable across engines and platforms such as Spark, Flink, Trino, Presto, Databricks, and Snowflake-adjacent environments
  • Debug performance bottlenecks across storage, metadata, table maintenance, query execution, network, and compute layers
  • Develop workload-aware table optimization systems that learn from access patterns and reorganize data automatically
  • Implement algorithms in compression, representation, layout optimization, and data efficiency
  • Contribute to open-source or publish research when appropriate

Requirements

  • Strong engineering depth in distributed systems, storage systems, databases, or data infrastructure
  • Production experience with modern data lake or lakehouse technologies such as Iceberg, Delta Lake, Hudi, Spark, Trino, Presto, Flink, Hive Metastore, Unity Catalog, or similar systems
  • Hands-on experience with columnar formats such as Parquet or ORC
  • Understanding of metadata-driven architectures, table formats, transaction semantics, query planning, and physical data layout
  • Experience with table maintenance, compaction, clustering, file sizing, metadata pruning, snapshot expiration, or garbage collection
  • Familiarity with cloud object storage systems such as S3, GCS, or ADLS and the performance tradeoffs of building lakehouse systems on top of them
  • Strong programming skills in Java, Scala, Go, Rust, C++, or similar systems-oriented languages
  • Curiosity about compression, entropy, information theory, and how data representation affects AI efficiency
  • A pragmatic builder’s mindset: rigorous, hands-on, and comfortable owning complex systems end to end

Qualifications

  • Experience contributing to Apache Iceberg, Delta Lake, Apache Hudi, Spark, Flink, Trino, Presto, Velox, DuckDB, Polars, Parquet, ORC, or related systems
  • Experience with manifests, snapshots, metadata catalogs, schema evolution, partition evolution, delete handling, transaction logs, or table garbage collection
  • Experience solving the small-file problem, optimizing object-store access patterns, or improving table health at scale
  • Background in storage engines, query engines, indexing, caching, encoding, compression, or adaptive query optimization
  • Research or open-source contributions in distributed systems, databases, storage, compression, indexing, or data processing
  • Interest in how physical data representation affects model training, inference, retrieval, and reasoning efficiency

Skills

  • Strong engineering depth in distributed systems, storage systems, databases, or data infrastructure
  • Production experience with modern data lake or lakehouse technologies such as Iceberg, Delta Lake, Hudi, Spark, Trino, Presto, Flink, Hive Metastore, Unity Catalog, or similar systems
  • Hands-on experience with columnar formats such as Parquet or ORC
  • Understanding of metadata-driven architectures, table formats, transaction semantics, query planning, and physical data layout
  • Experience with table maintenance, compaction, clustering, file sizing, metadata pruning, snapshot expiration, or garbage collection
  • Familiarity with cloud object storage systems such as S3, GCS, or ADLS and the performance tradeoffs of building lakehouse systems on top of them
  • Strong programming skills in Java, Scala, Go, Rust, C++, or similar systems-oriented languages
  • Curiosity about compression, entropy, information theory, and how data representation affects AI efficiency
  • A pragmatic builder’s mindset: rigorous, hands-on, and comfortable owning complex systems end to end

Benefits

  • Competitive salary, meaningful equity, and performance bonus for top performers
  • 401(k) with company match, comprehensive health coverage, and unlimited PTO
  • Daily catered meals in our Mountain View office
  • Support for research, publication, and conference participation

Pay

Compensation Range: $160K - $240K

Schedule

On-site

Similar jobs