Jobs · Engineering

Senior Site Reliability Engineer

Stord · United States · 6 days ago
RemoteRemoteEngineeringFull-time

About the role

The SRE team is small, fast-moving, and owns the infrastructure that keeps the platform running, primarily on Google Cloud Platform, across GKE, Cloud Run, AlloyDB, and the networking that connects our services. This is a high-autonomy environment with a wide surface area and few layers between you and the systems you're responsible for. The work is real infrastructure engineering: reliability, scale, cost, and the automation that lets a lean team punch well above its weight.

Responsibilities

  • Own architecture and implementation of scalable, reliable infrastructure on GCP, including GKE, Cloud Run, AlloyDB, and networking.
  • Own Infrastructure as Code in Terraform: modules, org policies, and the patterns the team builds on.
  • Manage containerized workloads on Kubernetes, including performance tuning, capacity planning, and resource optimization.
  • Drive down cost and toil through better defaults, right-sizing, and automation rather than manual intervention.
  • Build monitoring, alerting, and observability in Datadog (APM, logs, RUM) that catches problems before customers do.
  • Define the reliability signals that matter for the services you own, and hold the line on them.
  • Develop and maintain disaster recovery and business-continuity strategies, and prove they work.
  • Design and maintain CI/CD pipelines in GitHub Actions, including runner strategy and deployment safety.
  • Automate operational workflows and infrastructure provisioning so the platform scales smoothly as the team grows.
  • Build custom tooling and scripts that remove recurring operational pain.
  • Partner with data and development teams to improve deployment practices and application reliability.
  • Provide escalation support for production incidents, help lead post-incident reviews, and turn findings into durable fixes.
  • Help improve SRE and infrastructure best practices across the team, and participate in on-call for critical systems.

Requirements

Required 5+ years in SRE, platform, or infrastructure engineering: you've owned complex systems and driven technical work forward with minimal supervision. GCP depth: strong hands-on experience with GCP core services (GKE, Cloud Run, AlloyDB, networking, IAM). You know how these fit together in production, not just in a certification. Containers & orchestration: you're fluent in Docker and Kubernetes and can debug, tune, and scale real workloads. Infrastructure as Code: deep Terraform experience. You write reusable modules, reason about state, and treat infra changes with the same rigor as application code. A programming language you're genuinely productive in: TypeScript, Python, Go, or similar, used to build tooling and automation, not just glue scripts. Observability: you build monitoring and alerting that's actionable (Datadog, or equivalents like Prometheus/Grafana), and you know the difference between a noisy dashboard and a useful one. Distributed systems fundamentals: failure modes, consistency, and how systems break at scale. Git and collaborative development workflows: you work in shared codebases and review others' changes well. Incident management: you've run incidents and post-mortems and can stay calm and methodical when production is on fire. Required Soft Skills: Ownership & Accountability: You own features end-to-end and take pride in what you ship. You follow through from design to production and don't drop things. Strong Communication: You can explain technical decisions and trade-offs to engineers, PMs, and stakeholders. You ask good questions and listen well. Collaborative Approach: You work well with others, give constructive code review feedback, and actively seek input from teammates. Production Mindset: You prioritize reliability and user impact. You think about failure modes, monitoring, and operational concerns as part of your design process. Learning Agility: You're comfortable with rapidly evolving AI/ML technologies and tools. You stay current without chasing hype. Directed AI-Assisted Development: You know how to use AI coding tools as a productivity multiplier while maintaining quality and your own technical judgment.

Similar jobs