Jobs · Engineering

Senior Site Reliability Engineer (Cloud Platform)

Salve.Lab · Raleigh, NC · 2 wk ago
RemoteRemoteEngineeringFull-time

Key Responsibilities

  • Maintain the reliability, availability, and performance of production and pre-production environments.
  • Monitor platform health and improve alerting, automation, and operational processes.
  • Respond to production incidents, participate in root cause analysis, and implement long-term improvements.
  • Design, build, and enhance observability solutions using metrics, logs, traces, and dashboards.
  • Partner with software engineers to improve application reliability throughout the development lifecycle.
  • Develop and maintain operational documentation, troubleshooting guides, and runbooks.
  • Automate repetitive operational tasks to improve efficiency and reduce manual intervention.
  • Participate in on-call rotations while continuously improving incident response processes.
  • Promote reliability engineering principles, operational excellence, and continuous improvement across engineering teams.

Requirements

  • Bachelor's or Master's degree in Engineering, Computer Science, or a related field.
  • Strong experience operating Kubernetes or other container orchestration platforms.
  • Experience supporting large-scale production services.
  • Hands-on experience with AWS.
  • Experience with Prometheus, Grafana, and ELK.
  • Strong scripting skills (Bash, Python, or Go).
  • Experience administering Linux-based production environments.
  • Experience with Infrastructure as Code or configuration management tools such as Terraform or Ansible.
  • Solid understanding of networking fundamentals (TCP/IP, DNS, load balancing, routing).
  • Excellent troubleshooting, communication, and collaboration skills.
  • A proactive mindset with a passion for automation and reliability.

Nice to Have

  • Experience with SIP or VoIP technologies.
  • Familiarity with MySQL or PostgreSQL.
  • Experience with Redis or other NoSQL databases.

Similar jobs