Jobs · Engineering · California

Site Reliability Engineer

Aurora · Palo Alto, CA · Yesterday
HybridEngineering$170k–$250k/yrFull-time

About the role

This is a site reliability role for someone who wants ownership over the systems that keep a multi-cloud GPU orchestration platform dependable in production. You will own observability, SLOs, automation, infrastructure-as-code, and incident response for services that sit close to the customer experience and the company's revenue. The goal is to convert reliability into engineering primitives — SLIs, SLOs, automation, and safe deployment patterns — so operational burden does not scale linearly with demand. You will work directly with the founding engineering team and help define how infrastructure engineering operates as the company scales.

The technical problem

GPU infrastructure has failure modes that are expensive and messy: cloud-provider drift, partial outages, capacity swings, noisy neighbors, inconsistent health signals, and operational work that grows quickly if it is not automated early. The hard part is not adding more dashboards. The hard part is reconciling provider state with internal orchestration logic, keeping visibility high across clouds, and building guardrails that let the team move quickly without accumulating hidden reliability debt.

What you'll own

  • Observability: build and own dashboards, alerts, and distributed tracing with OpenTelemetry, Prometheus, and Grafana so the team can see system behavior at the level that matters.
  • SLIs and SLOs: define and implement reliability targets across the API layer and internal orchestration services, and make sure new features are designed with operability in mind.
  • Automation: write Python or Go to remove repetitive operational work, including provider API reconciliation, automated health checks, and capacity rebalancing.
  • Infrastructure-as-code: maintain and extend Terraform or Pulumi modules and Kubernetes configurations as the multi-cloud footprint expands.
  • Incident response: participate in on-call, drive root-cause analysis, and convert incidents into durable fixes instead of recurring pages.
  • Operational playbook: help shape the standards for rollouts, monitoring, runbooks, and reliability reviews so the infrastructure surface stays manageable as the company grows.

Who this is for

You are likely a strong fit if you have:

  • 3+ years in SRE, Production Engineering, or Infrastructure roles with real production ownership.
  • Built automation, observability, or platform tooling end to end, not just contributed tickets inside a larger system.
  • Worked in environments where distributed systems behavior, failure handling, and operational hygiene were core to the job.
  • Comfort across multi-cloud infrastructure and the failure modes that come with provider APIs, orchestration, and partial outages.
  • Strong instincts for Kubernetes, Terraform or Pulumi, and operationally safe rollouts.
  • Experience with Python or Go for infrastructure automation.
  • Ability to define useful SLIs and SLOs and reason about them in production, not just in design docs.
  • The judgment to improve reliability without creating a heavier process layer than the team needs.
  • Bonus: exposure to GPU, AI/ML, or other accelerated compute workloads.

Tech stack

  • Automation: Python, Go
  • Infrastructure: Terraform, Pulumi, Kubernetes
  • Observability: OpenTelemetry, Prometheus, Grafana
  • Environment: multi-cloud GPU infrastructure and orchestration services

You do not need to be deep in every tool above, but you should understand how they fit together in a production system with expensive compute and tight reliability constraints.

Why now

The company has already crossed the line from early validation to real usage: growing revenue, real customers, and enough scale that manual operational work will become a drag unless it is engineered away. This is the point where reliability work stops being a support function and becomes part of the product architecture. The engineer in this seat will help set the standard for how the platform is observed, operated, and debugged for the next stage of growth.

This role is not for you if

  • You want a narrowly scoped role with fully specified tickets.
  • You prefer owning alerts over owning the systems that generate them.
  • You are uncomfortable with on-call and incident accountability.
  • You want to avoid infrastructure work that includes automation, rollouts, and systems design.
  • You need a mature playbook before you can make judgment calls.

Compensation and logistics

  • Base salary: $170K–$250K
  • Equity: competitive
  • Employment: full-time
  • Location: San Francisco or Palo Alto, CA
  • Work model: on-site 4 days per week
  • Cadence: all team members come to the Palo Alto office on Mondays
  • Visa support: transfer-only sponsorship support available

Similar jobs

Site Reliability Engineer

QualityAIIllinois, United States· Today
Quality Assurance$110k–$120k/yrapply on careers.quality-ai.com