Jobs · Engineering · Washington

Senior Software Engineer - Capacity

Snowflake · Bellevue, WA · 6 days ago
Engineering$200k–$288k/yrFull-time

About the role

At Snowflake, we are powering the era of the agentic enterprise. We seek AI-native thinkers who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic, fast-moving environments and approach challenges with an experimental mindset, rapidly testing emerging capabilities to discover simpler, more powerful ways to deliver results.

Snowflake’s infrastructure is expanding rapidly across AWS, Azure, and GCP. The Capacity team plays a pivotal role in provisioning the cloud resources essential for Snowflake's operations and ongoing growth. Capacity Engineering accurately models demand, forecasts requirements, and delivers optimal CPU and GPU capacity on schedule. We drive hardware cost-efficiency and price/performance while continually maximizing fleet utilization.

To achieve this across all major cloud providers, the team is building a centralized, self-serve internal capacity platform. This software-driven system provides early visibility into supply risks, ensures sufficient lead time for capacity deployment, and maintains high utilization across committed cloud resources. The technical problem spans the full lifecycle: modeling demand and supply as first-class data, reconciling heterogeneous provider telemetry and commitments into a single canonical capacity layer, and making the live state of the fleet legible and actionable in real time.

We are actively looking for a senior software engineer who loves solving problems at scale, prefers to write scalable, reliable, and testable software, is an ace troubleshooter, and is deeply technical. Snowflake’s growth and multi-cloud footprint in a constrained capacity environment demand real engineering maturity in the systems that plan and land compute.

Responsibilities

  • Design and build the capacity platform that unifies CPU and GPU allocation, procurement, reservation lifecycle, and utilization across all three clouds.
  • Own the canonical capacity data layer: ingest and reconcile demand forecasts, provider supply signals, commitments, and fleet utilization into a single, trustworthy model consumed across the company.
  • Serve as the liaison to Cloud Service Providers managing and integrating vendor relationships into the capacity planning and procurement workflow.
  • Build planning and allocation systems that translate demand into hardware requirements (shape, quantity, region, timing) and surface supply risk early, with real-time visibility into fleet and reservation health.
  • Drive efficiency: instrument utilization across CPU and GPU accelerator workloads, establish price/performance baselines, and build the tooling that recovers stranded capacity and right-sizes commitments.
  • Integrate hardware evolution into the platform: evaluate new CPU and GPU generations and their price/performance, and build the flexibility (backup and cross-family fallbacks) that keeps plans aligned to the hardware roadmap.
  • Partner with core services, warehouse, AI/ML, and finance teams to forecast and procure capacity ahead of launches, support AI/ML workloads reliably, and turn insights into procurement and allocation decisions.
  • Ensure high availability, reliability, and performance of capacity systems by participating in on-call rotations and incident management.

Requirements

  • 7+ years of industry experience designing, building, and supporting large-scale systems in production.
  • Hands-on experience working with cloud providers on compute cluster and cloud services provisioning (CPU and/or GPU fleets).
  • Experience with capacity planning, procurement, resource management, or efficiency work on systems built on large private clouds or public cloud providers.
  • Deep system and architectural analysis experience to identify actionable performance, availability, and efficiency insights across CPU and GPU accelerator fleets.
  • Proficiency in programming languages such as Go, Python, or Java.
  • Excellent problem-solving skills and ability to troubleshoot complex issues in a production environment.
  • Strong communication skills and the ability to collaborate effectively in a team environment.
  • BS / MS in Computer Science, Engineering or related fields.

Nice to have

  • Experience developing or using observability infrastructure such as OpenTelemetry or Prometheus.
  • Familiarity with accelerator/GPU fleets, hardware price/performance analysis, or Kubernetes-based compute at scale.
  • Prior background working with Modeling, Forecasting, Cloud Spend Optimization, and LLMs.

Pay

The estimated base salary range for this role is $200,000 - $287,500. Additionally, this role is eligible to participate in Snowflake’s bonus and equity plan.

Benefits

  • Medical, dental, vision, life, and disability insurance.
  • 401(k) retirement plan.
  • Flexible spending & health savings account.
  • At least 12 paid holidays.
  • Paid time off.
  • Parental leave.
  • Employee assistance program.
  • Other company benefits.

Similar jobs

Senior Software Engineer

PIADA ITALIAN STREET FOODColumbus, OH· 2 days ago
$120k–$180k/yrapply on careers-thepiadagroup.icims.com

Senior Software Engineer

Delta Dental Ins.Alpharetta, GA· 2 mo ago
Information Technology$152k–$154k/yrapply on ejep.fa.us2.oraclecloud.com

Senior Software Engineer

Johns Hopkins Applied Physics LaboratoryLaurel, MD· 1 mo ago
Engineering$105k/yrapply on careers.jhuapl.edu