Site Reliability Analyst
Haystack · St Louis, MO · 2 wk ago
RemoteRemoteManagementFull-time
About the role
Drive end-to-end reliability, performance, observability, and operational excellence for a large-scale enterprise performance analytics platform. Initially support system maintenance and operations, transitioning into a performance engineering focus. Support and maintain the Capacity Management (CPAC) environment, including daily operational monitoring and health checks. Troubleshoot complex performance issues, collector failures, missing data, and application stability concerns. Design, build, and support a production monitoring and analytics platform in a Docker/Kubernetes containerized environment. Develop operational runbooks, standard operating procedures, and support documentation.
Requirements
- 5+ years of Linux/Unix administration experience, including system build and deployment.
- 5+ years of experience supporting enterprise production applications.
- 3+ years of container platform administration experience (Docker/Kubernetes).
- Proven experience troubleshooting performance issues in distributed production environments and performing root cause analysis.
- Programming and automation experience with Python, NoSQL, Svelte, SQL, GIT, or Bash/Shell scripting.
- Experience with observability tools such as Splunk, Grafana, Prometheus, AppDynamics, or OpenTelemetry.
Benefits
- Comprehensive benefits package, including medical, dental, vision, life, and disability.
- Employee Stock Purchase Program (ESPP) and 401K program with company match.
- Access to on-demand training, certification prep, and a library of technical and leadership courses.
- Dedicated customer service team and certified Career Coach support.