Jobs · Engineering

Staff Site Reliability Engineer

RemoteHunter · United States · 1 wk ago
RemoteRemoteEngineering$180k–$240k/yrFull-time

Our client operates in the AI marketing platform sector, focusing on 1:1 personalization to strengthen brand and customer connections. They provide a unified platform integrating SMS, RCS, email, and push notifications, using AI and real-time behavioral data to deliver personalized customer experiences. Serving more than 8,000 customers across over 70 industries, the organization facilitates billions of customer interactions and supports leading global brands with a distributed workforce.

About the Role

The Staff Site Reliability Engineer plays a strategic role in improving the reliability, scalability, and performance of the organization’s platform infrastructure. This position designs and implements solutions that strengthen observability, traceability, incident management, and platform scalability. The role also provides technical leadership across teams, mentors engineers, establishes production standards, and influences the technical roadmap to help engineering teams deliver reliable and secure solutions efficiently.

Responsibilities

  • Design and implement systems that improve reliability, observability, traceability, and incident management
  • Lead strategic cross-team projects and provide technical leadership
  • Collaborate with AI/ML, Data, Platform, and Product teams to develop advanced services
  • Define and enforce production standards, processes, and tools
  • Establish and implement reliability metrics, including SLIs and SLOs
  • Mentor and guide team members to support technical growth and development
  • Drive continuous improvement by introducing innovative solutions and challenging existing practices

Requirements

  • 7+ years of experience in Production Engineering, Backend Engineering, SRE, DevOps, or a related field
  • Strong technical vision and ability to plan for future platform needs
  • Proficiency in at least one programming language, such as Golang, Python, Java, or TypeScript
  • Proven experience delivering medium- to large-scale projects that improve platform reliability and scalability
  • Deep understanding of production reliability concepts, including SLIs, SLOs, and incident management
  • Excellent communication skills with the ability to collaborate across technical and non-technical teams
  • Experience working in dynamic, reliability-focused production environments preferred

Pay

US base salary range of $180,000 to $240,000 annually, plus equity compensation and benefits. Compensation may vary based on role, level, location, and other relevant factors.

Benefits & Perks

  • Health and wellness benefits
  • Equity compensation

Similar jobs