Jobs · Information Technology · California

Senior Software Engineer — Distributed Compute / Spark Systems

Granica · San Francisco Bay Area · Yesterday
HybridInformation Technology$160k–$240k/yrFull-time

About the role

Granica is hiring a Senior Software Engineer to build distributed compute systems for enterprise-scale data and AI workloads. This includes systems for distributed execution, workload optimization, query performance, scheduling, resource management, and compute cost reduction across petabyte- and exabyte-scale environments.

You will work on the core infrastructure behind Crunch, Granica’s continuous optimization product for enterprise lakehouse data. This role involves building systems that directly affect customer compute spend, query latency, workload reliability, cluster efficiency, and the performance of large-scale analytical data processing.

What You’ll Do

  • Build distributed compute systems for large-scale analytical and AI workloads
  • Improve performance and cost efficiency across Spark, Trino, Presto, Flink, Databricks, and Snowflake-adjacent environments
  • Design workload-aware systems for query execution, resource allocation, scheduling, and compute optimization
  • Optimize execution performance across joins, aggregations, scans, shuffles, spills, caching, partitioning, and task scheduling
  • Build systems that learn from workload patterns and automatically improve execution plans, cluster usage, and compute efficiency
  • Develop infrastructure for adaptive workload routing, execution planning, and data-processing reliability across large customer environments
  • Debug performance bottlenecks across query execution, metadata, storage, network, memory, CPU, and distributed compute layers
  • Build systems that reduce compute waste caused by inefficient scans, poor partitioning, small files, skew, unnecessary shuffles, and suboptimal workload placement
  • Improve reliability and failure recovery for large distributed data-processing jobs
  • Implement algorithms in workload optimization, execution efficiency, cost modeling, and data-processing performance
  • Contribute to open-source or publish research when appropriate

What We’re Looking For

  • Strong engineering depth in distributed systems, data processing systems, query engines, databases, or cloud infrastructure
  • Production experience with distributed compute or query systems such as Apache Spark, Spark SQL, Trino, Presto, Flink, Databricks, EMR, Glue, Hive, or similar systems
  • Hands-on experience improving performance, reliability, or cost efficiency for large-scale data-processing workloads
  • Understanding of distributed execution, query planning, scheduling, resource management, fault tolerance, and workload isolation
  • Experience with Spark internals, Spark SQL, Catalyst, Adaptive Query Execution, shuffle, joins, aggregation, spill, memory management, or task scheduling
  • Familiarity with lakehouse formats and columnar data such as Iceberg, Delta Lake, Hudi, Parquet, or ORC
  • Familiarity with cloud object storage systems such as S3, GCS, or ADLS and the performance tradeoffs of running distributed compute on top of them
  • Strong programming skills in Scala, Java, Go, Rust, C++, or similar systems
  • Curiosity about workload optimization, cost modeling, adaptive execution, and how compute efficiency affects AI and analytics at scale
  • A pragmatic builder’s mindset: rigorous, hands-on, and comfortable owning complex systems end to end

Bonus Experience

  • Contributing to Apache Spark, Spark SQL, Trino, Presto, Flink, Velox, DuckDB, DataFusion, Iceberg, Delta Lake, Hudi, Parquet, ORC, or related systems
  • Experience with Catalyst, Adaptive Query Execution, cost-based optimization, query planning, vectorized execution, or distributed runtime systems
  • Experience optimizing joins, aggregations, shuffles, scans, spills, caching, partitioning, skew handling, or task scheduling
  • Experience building workload schedulers, execution control planes, resource managers, or multi-engine compute platforms
  • Experience reducing compute cost or improving workload efficiency in large-scale production data environments

Why Join Granica

  • Build foundational infrastructure for enterprise data and AI
  • Work on deep systems problems across distributed compute, query execution, workload optimization, scheduling, resource management, and compute efficiency
  • Partner directly with Product, Engineering, and company leadership
  • Help shape Crunch, Granica’s production data optimization platform for enterprise-scale lakehouse environments
  • Work with a small, high-caliber team solving high-value infrastructure problems at massive scale
  • Have direct influence on architecture, product direction, customer outcomes, and company growth

Compensation & Benefits

  • Competitive salary, meaningful equity, and performance bonus for top performers
  • 401(k) with company match, comprehensive health coverage, and unlimited PTO
  • Daily catered meals in our Mountain View office
  • Support for research, publication, and conference participation

At Granica

You'll help build the next generation of enterprise AI—from exabyte-scale data infrastructure, Large Tabular Models (LTMs), and stateful AI agents. Together, we're creating the infrastructure that enables enterprises to own their data, own the intelligence built on it, and scale both efficiently.

Similar jobs