Jobs · Engineering · California

Site Reliability Engineer

Forward · Santa Clara, CA · 6 days ago
Engineering$230k–$250k/yrFull-time

About the role

This is not a "keep the lights on" SRE role. As our first or early SRE hire you will be building the reliability engineering function at Forward — defining how we think about availability, observability, incident response, and operational excellence across a complex, distributed SaaS platform. You will work closely with engineering, infrastructure, and product to ensure our platform meets the reliability bar our enterprise customers demand.

Responsibilities

  • Define and drive SRE practices from the ground up — SLOs, SLIs, error budgets, and the frameworks the engineering org will actually use
  • Drive the reliability and operational excellence of the Forward SaaS platform
  • Build and maintain observability infrastructure — logging, metrics, tracing, and alerting — so the team always knows what's happening before customers do
  • Lead incident response: on-call rotations, runbooks, post-mortems, and the follow-through to make sure the same incident doesn't happen twice
  • Partner with engineering teams to embed reliability thinking into the SDLC — capacity planning, load testing, chaos engineering, and production readiness reviews
  • Help define and build the SRE team as the company scales — this is a foundational hire with a path to leadership

Requirements

  • 6+ years of experience in site reliability engineering, DevOps, or infrastructure engineering in a SaaS or cloud environment
  • Proven experience building or significantly maturing an SRE function — not just operating within one someone else built
  • Strong fundamentals in networking — TCP/IP, DNS, routing, switching, firewalls, and load balancing. Experience with network management or observability platforms is a significant plus
  • Hands-on experience with Kubernetes and container orchestration in production environments
  • Deep proficiency with observability tooling — Prometheus, Grafana, Datadog, Splunk, or similar
  • Strong scripting and automation skills in Python, Bash, or similar
  • Experience with cloud platforms — AWS, GCP, or Azure — including infrastructure as code (Terraform, Ansible, or equivalent)
  • Track record of owning and improving incident response processes including blameless post-mortems and SLO-driven reliability improvements
  • Able to communicate clearly with both engineering teams and non-technical stakeholders — you can explain an outage to a customer-facing team without jargon and explain an SLO to an executive without losing them

Qualifications

  • Nice to have experience supporting enterprise or federal government customers with high availability requirements
  • Experience in a foundational or early SRE hire capacity at a growth stage company

Skills

  • Networking fundamentals
  • Kubernetes and container orchestration
  • Observability tooling
  • Cloud platforms
  • Communication and stakeholder management

Benefits

  • Competitive compensation, equity, and the opportunity to grow into a leadership role as the SRE function scales

Pay

The base pay range for this role is between $230,000 and $250,000. Base pay will depend on your skills, qualifications, experience, and location.

Schedule

Not specified

Similar jobs

Site Reliability Engineer

Akamai TechnologiesCambridge, MA· 5 days ago
RemoteEngineering$76k–$136k/yrapply on fa-extu-saasfaprod1.fa.ocs.oraclecloud.com