Software Engineer – Presto / Spark Development
IBM · San Jose, CA · 5 days ago
HybridEngineeringFull-time
Your Role And Responsibilities
- Develop Engine Internals: Build and maintain Presto and/or Spark connectors, operator implementations, optimizer rules, and custom UDFs/UDAFs for open table formats (Iceberg, Delta, Hudi).
- Optimize Performance at Scale: Diagnose and resolve data skew, broadcast-join sizing, shuffle bottlenecks, and memory pressure; tune operator memory, spill thresholds, and off-heap usage using async-profiler and flamegraphs.
- Contribute to CI/CD & Benchmarks: Contribute to the automated CI/CD pipeline and maintain CI benchmarks that guard against regressions in query latency, throughput, and resource consumption.
- Support Production & Debug: Support Presto/Spark deployments on Kubernetes and bare metal, unit-test fixes for engine-related and customer-reported issues, and participate in on-call.
- Collaborate in Agile Environment: Partner with query optimization, storage, GPU acceleration, and AI/ML teams, conduct reviews with measurable acceptance criteria, and document connector interfaces and engine internals.
Preferred Education
Bachelor's Degree Required
Technical And Professional Expertise
- Engine Development Experience: 6+ years of professional software engineering, including at least 2 years developing against Presto/Trino or Apache Spark internals.
- Language & Codebase Proficiency: Strong Java or Scala skills with comfort navigating and modifying a large, complex open-source codebase.
- Connectors & Optimizer Work: Hands-on experience building connectors, UDFs/UDAFs, or optimizer extensions, plus working knowledge of Spark/Presto query planning, physical execution, and the operator/stage model.
- Performance & Formats: Experience resolving data skew, shuffle bottlenecks, broadcast-join sizing, and memory pressure at scale; familiarity with an open table format (Iceberg, Delta, or Hudi); JVM tuning in production.
- Communication & Education: Clear written communication—able to file actionable bugs, write design docs, and explain engine trade-offs; Bachelor's degree in Computer Science, Engineering, or equivalent practical experience.
Preferred Technical And Professional Experience
- OSS & GPU Acceleration: Committer or significant contributor to Apache Spark, Trino, or Presto, and experience integrating GPU-accelerated execution (RAPIDS Accelerator, cuDF) into query paths.
- Vectorization & Multi-Tenancy: Familiarity with vectorized execution and columnar formats (Arrow, ORC, Parquet), ML feature and inference pipelines on Spark/Presto, FinOps cost modeling, and multi-tenant deployments with fairness scheduling and workload isolation.