Jobs · Engineering · California

Software Engineer, Spark Platform

DoorDash · San Francisco, CA · 1 mo ago
Engineering$131k–$192k/yrFull-time

About The Team

The Spark Platform team owns and operates DoorDash's Apache Spark ecosystem — the execution runtime, remote shuffle service, cluster scheduler, and reliability tooling that powers the company's data, analytics, and ML workloads. We run Spark across the company at significant scale and continue to expand the workloads, capabilities, and consumer base we serve. Orchestrating and operating thousands of Spark cluster deployments is a complex distributed system problem which the team invests heavily in runtime optimization, systems architecture, multi-tenant scheduling, and end-user tooling.

About The Role

As a Software Engineer on Spark Platform, you will execute across the surfaces of our in-house Spark deployment that serves the entire company. The work spans Spark runtime upgrades and performance, multi-tenant scheduling and executor bin-packing on Kubernetes, cluster lifecycle automation, and the observability and incident automation that keep the platform sustainable. You will move between layers as the work demands — picking up the next high-leverage problem regardless of where it sits — and partner closely with the rest of the team and with platform consumers across the company.

Requirements

  • You are located or willing to relocate to the Bay Area, Seattle, or NYC.
  • B.S., M.S., or PhD in Computer Science or equivalent.
  • 24+ years of industry experience operating production distributed systems.
  • Experience operating Apache Spark at scale on Amazon EMR, Databricks, or an in-house deployment — with a focus on platform operations (runtime upgrades, cluster lifecycle, shuffle, observability, multi-tenant scheduling) rather than authoring individual Spark jobs.
  • Hands-on experience operating production systems on Kubernetes — controllers, operators, custom resources, and the failure modes that show up in multi-tenant clusters.
  • Familiarity with batch or big-data schedulers (YuniKorn, Volcano, Kueue, or equivalent) and/or with the Spark-on-Kubernetes operator.
  • Familiarity with observability stacks (Prometheus, OpenTelemetry, distributed tracing, structured logging) and with defining SLOs and SLIs that change team behavior.
  • Comfort working in a cloud environment (AWS preferred) — VPC networking, instance lifecycle, spot/preemptible markets, and autoscaling primitives.
  • Professional experience with Python, Go, Scala, or Java; SQL fluency.
  • A bias toward incremental rollout, measurement, and reducing toil.

Qualifications

  • Experience with Python, Go, Scala, or Java.
  • SQL fluency.

Skills

  • Operate production distributed systems.
  • Operate Apache Spark at scale.
  • Operate production systems on Kubernetes.
  • Operate batch or big-data schedulers.
  • Operate observability stacks.
  • Define SLOs and SLIs.
  • Work in a cloud environment.
  • Use Python, Go, Scala, or Java.
  • Use SQL.

Benefits

We offer a comprehensive benefits package to all regular employees, including:

  • A 401(k) plan with employer matching.
  • 16 weeks of paid parental leave.
  • Wellness benefits.
  • Commuter benefits match.
  • Paid time off and paid sick leave in compliance with applicable laws (e.g. Colorado Healthy Families and Workplaces Act).
  • Medical, dental, and vision benefits.
  • 11 paid holidays.
  • Disability and basic life insurance.
  • Family-forming assistance.
  • A mental health program.

Pay

The national base pay ranges for this position within the United States, including Illinois and Colorado.

  • I4: $130,600—$192,000 USD
  • I5: $159,800—$235,000 USD
  • I6: $193,800—$285,000 USD

Schedule

This is a hybrid position.

Similar jobs