Jobs · Engineering · Georgia

Site Reliability Engineer

Incident IQ · Alpharetta, GA · 3 days ago
EngineeringFull-time

Site Reliability Engineer (SRE)

This role is hands-on from day one. Your initial focus will be:

  • SLI/SLO Definition & Grafana Implementation: Drive the definition of Service Level Indicators (SLIs) and Service Level Objectives (SLOs) for our core services, translating them into insightful Grafana dashboards and actionable, burn-rate-based alerting, so pages are precise and noise stays low.
  • Incident Management: Stand up our incident management practice (tooling such as PagerDuty, on-call training, incident command), then own and continuously improve it, stepping in personally only for the most severe incidents.
  • Observability Stack Ownership: Own the observability stack end to end: metrics, logs, traces, Real User Monitoring (RUM), and synthetic checks across the user journey, alerting whenever a signal deviates from baseline.
  • Team Enablement: Partner with engineering teams to refine SLIs, SLOs, and error budgets as services evolve, and coach teams on SRE and observability best practices.
  • Toil Reduction: Identify and automate away manual, repetitive operational work through infrastructure as code and tooling.
  • Chaos & Performance Engineering: Design and run load/performance tests and chaos engineering game days to proactively surface weaknesses before they cause incidents.

Requirements

  • Education & Systems Foundations: Bachelor's degree in Computer Science, Computer Engineering, or equivalent formal training, with real depth in operating systems, databases, and networking. This fundamental understanding is required.
  • AI-Accelerated Execution (core requirement): You actively use AI tools daily to multiply your own output, not just experiment with them on the side. We expect you to use AI to write and debug code faster, stand up dashboards and alerts faster, and generally ship at a pace that wouldn't be possible without it.
  • Track Record Over Tenure: A demonstrated history of independently driving big, ambiguous reliability or infrastructure projects to completion, typically reflecting 5+ years in an SRE, DevOps, or production engineering role. We care far more about what you've actually shipped than the number itself.
  • SLI/SLO Methodology: Proven, hands-on track record implementing the SLI/SLO/error-budget model in a prior role, the discipline formalized in Google's SRE Workbook.
  • Observability Tooling: Strong experience with Grafana and PromQL (Prometheus Query Language), Grafana Alloy for Loki logs, and a metrics backend such as Prometheus or Datadog. Experience instrumenting with OpenTelemetry and a tracing/Application Performance Monitoring (APM) backend (open-source preferred: SigNoz, Uptrace, Tempo; commercial: Datadog, New Relic), plus Real User Monitoring (RUM) and synthetic monitoring (e.g., Grafana Faro, Grafana Synthetic Monitoring / k6).
  • Incident Management: Proven track record designing on-call rotations and incident command practices elsewhere, with tooling such as PagerDuty or equivalent.
  • Performance & Chaos Engineering: Hands-on with a load/performance framework (Locust, k6, or JMeter) and chaos engineering exercises to validate reliability under real conditions.
  • Automation, Infrastructure & Cloud: Proficient in Python, Go, or Bash; hands-on with Infrastructure as Code (Terraform, Ansible, or equivalent), Kubernetes, and at least one major cloud platform (Amazon Web Services (AWS), Google Cloud Platform (GCP), or Azure).
  • Communication: Experienced, versatile communicator: able to go deep with developers on root cause, tradeoffs, and implementation detail; comfortable pushing back with a real technical path when a team says something “can't” be done; precise about the difference between a mitigation and an actual fix when reporting status; and able to translate reliability status, risk, and priorities clearly for business and engineering stakeholders.
  • Independence & Pace: You don't need hand-holding or check-ins to make progress. Comfortable resolving ambiguous problems in hours, not weeks, taking full ownership of outcomes, and juggling multiple threads under real time pressure without dropping the ball.

Similar jobs

Site Reliability Engineer

Akamai TechnologiesCambridge, MA· 5 days ago
RemoteEngineering$76k–$136k/yrapply on fa-extu-saasfaprod1.fa.ocs.oraclecloud.com