#23183 - Site Reliability Engineer
Qualitest · Chicago, IL · 1 mo ago
HybridConsulting$110k–$120k/yrFull-time
About the role
We are seeking an experienced Site Reliability Engineer (SRE) with a strong background in AWS Cloud, monitoring/observability platforms, and automation. The ideal candidate will partner closely with application development teams to improve application reliability, resiliency, performance, and operational excellence across hybrid cloud environments.
Responsibilities
- Partner with application development teams to improve application resiliency and reliability.
- Implement and maintain Service Level Objectives (SLOs), Service Level Indicators (SLIs), and operational best practices.
- Build end-to-end observability solutions using monitoring, logging, and tracing tools.
- Design and implement monitoring, alerting, dashboards, and health checks for production applications.
- Automate operational processes to reduce manual effort and improve system reliability.
- Develop and enhance capacity planning and performance management capabilities.
- Support disaster recovery (DR) planning and implementation for critical applications.
- Develop and support chaos engineering and resilience testing initiatives.
- Participate in production incident management and on-call support rotation.
- Collaborate with cross-functional teams to improve system availability, scalability, and operational efficiency.
Requirements
- 6–12 years of professional experience as a Site Reliability Engineer (SRE)
- Strong hands-on experience with AWS Cloud applications and services (mandatory)
- Experience supporting hybrid environments (AWS Cloud and on-premises deployments)
- Strong Linux/Unix administration and shell scripting experience
- Experience with Systems Observability and Application Performance Monitoring (APM) tools, preferably Datadog (Dynatrace experience is also valuable)
- Experience building dashboards using Grafana and Kibana
- Experience in performance testing and the ability to translate functional and non-functional requirements into automated non-functional testing (NFT) solutions
- Strong understanding of application integration, high availability, resilience, and observability
- Experience with DevOps practices and CI/CD pipelines
- Strong programming skills in one or more of the following: Python, Java, Shell Scripting (Unix/Linux)
- Strong understanding of Software Development Life Cycle (SDLC)
Qualifications
- Hands-on experience with ServiceNow (SNOW)
- Experience with container technologies such as Kubernetes and OpenShift
- Experience with Jenkins and CI/CD automation
- Experience with automation tools such as Ansible
- Strong working knowledge of JIRA
- Basic understanding of Release Management
- Understanding of Agile methodologies
- Experience with AWS Lambda services
- Knowledge of SQL, MySQL, and database concepts
Skills
- Cloud: AWS (mandatory), Hybrid Cloud, On-Prem, Linux/Unix, AWS Lambda
- Mandatory: Monitoring & Observability: Datadog (preferred), Dynatrace, Grafana, Kibana, ELK, APM tools
- Programming: Python, Java, Shell Scripting, Go (preferred), Ansible
- DevOps: Jenkins, CI/CD, Kubernetes, OpenShift
Benefits
- Competitive pay, the salary range for the role is $110,000 - $120,000
Pay
- $110,000 - $120,000
Schedule
- Hybrid – 2 to 3 days/week onsite