Jobs · Engineering · California

Senior Software Engineer — Lakehouse Systems

Agilesoft · San Francisco Bay Area · Yesterday
On-siteEngineering$160k–$240k/yrFull-time

About the role

Granica builds AI infrastructure for enterprises operating massive data environments. This role involves building foundational lakehouse systems for AI.

Responsibilities

  • Build metadata and transaction systems for large-scale tabular datasets
  • Design systems that support time travel, schema evolution, partition evolution, snapshot isolation, and atomic consistency
  • Develop table-maintenance infrastructure for lakehouse formats such as Apache Iceberg, Delta Lake, and Apache Hudi
  • Build systems for manifests, snapshots, transaction logs, metadata pruning, snapshot expiration, table garbage collection, and catalog consistency
  • Optimize file layout, clustering, compaction, file sizing, data skipping, indexing, and read-path performance
  • Improve performance and cost efficiency across object-store-backed lakehouse environments such as S3, GCS, and ADLS
  • Work with columnar formats such as Parquet and ORC, including encoding, compression, layout, pruning, and read-path optimization
  • Build systems that make lakehouse tables faster, cheaper, and more reliable across engines and platforms such as Spark, Flink, Trino, Presto, Databricks, and Snowflake-adjacent environments
  • Debug performance bottlenecks across storage, metadata, table maintenance, query execution, network, and compute layers
  • Develop workload-aware table optimization systems that learn from access patterns and reorganize data automatically
  • Implement algorithms in compression, representation, layout optimization, and data efficiency
  • Contribute to open-source or publish research when appropriate

Requirements

Strong engineering depth in distributed systems, storage systems, databases, or data infrastructure

Production experience with modern data lake or lakehouse technologies such as Iceberg, Delta Lake, Hudi, Spark, Trino, Presto, Flink, Hive Metastore, Unity Catalog, or similar systems

Hands-on experience with columnar formats such as Parquet or ORC

Understanding of metadata-driven architectures, table formats, transaction semantics, query planning, and physical data layout

Experience with table maintenance, compaction, clustering, file sizing, metadata pruning, snapshot expiration, or garbage collection

Familiarity with cloud object storage systems such as S3, GCS, or ADLS and the performance tradeoffs of building lakehouse systems on top of them

Strong programming skills in Java, Scala, Go, Rust, C++, or similar systems-oriented languages

Curiosity about compression, entropy, information theory, and how data representation affects AI efficiency

A pragmatic builder’s mindset: rigorous, hands-on, and comfortable owning complex systems end to end

Qualifications

Experience contributing to Apache Iceberg, Delta Lake, Apache Hudi, Spark, Flink, Trino, Presto, Velox, DuckDB, Polars, Parquet, ORC, or related systems

Experience with manifests, snapshots, metadata catalogs, schema evolution, partition evolution, delete handling, transaction logs, or table garbage collection

Experience solving the small-file problem, optimizing object-store access patterns, or improving table health at scale

Skills

Storage engines, query engines, indexing, caching, encoding, compression, or adaptive query optimization

Benefits

Competitive salary, meaningful equity, and performance bonus for top performers

401(k) with company match, comprehensive health coverage, and unlimited PTO

Daily catered meals in our Mountain View office

Support for research, publication, and conference participation

Similar jobs