Site Reliability Engineer
About the role
Careerscape is supporting a client opening for a Hybrid Site Reliability Engineer. This role focuses on maintaining highly available systems, improving infrastructure reliability, automating operational processes, monitoring production environments, and ensuring scalable cloud-based services operate efficiently in a hybrid work environment. This is an excellent opportunity for someone with strong technical skills who is passionate about automation, cloud technologies, system reliability, and continuous improvement. The Site Reliability Engineer will work closely with software engineers, DevOps teams, security professionals, and infrastructure teams to optimize system performance, resolve production issues, improve deployment pipelines, and build reliable, scalable platforms. This role is ideal for candidates interested in cloud infrastructure, DevOps engineering, platform engineering, or production operations.
Compensation
The salary range for this role is $105,000 - $115,000 per year, based on experience, technical expertise, certifications, and client hiring needs. Additional compensation and benefits details will be shared during the hiring process.
What you'll do
- Monitor production systems and cloud infrastructure
- Improve application availability, reliability, and scalability
- Build and maintain infrastructure automation solutions
- Manage CI/CD pipelines and deployment processes
- Respond to incidents and perform root cause analysis
- Optimize system performance and resource utilization
- Develop monitoring, logging, and alerting solutions
- Collaborate with engineering teams to improve system architecture
- Maintain infrastructure as code using automation tools
- Support disaster recovery and business continuity initiatives
- Implement security and reliability best practices
- Perform additional infrastructure and platform engineering duties as assigned
What we're looking for
- Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field preferred
- 2–5 years of experience in Site Reliability Engineering, DevOps, Systems Engineering, Cloud Infrastructure, or related roles
- Strong knowledge of Linux system administration
- Experience with AWS, Microsoft Azure, or Google Cloud Platform
- Experience with Kubernetes, Docker, and container orchestration
- Proficiency with Infrastructure as Code tools such as Terraform or CloudFormation
- Experience with CI/CD tools including GitHub Actions, Jenkins, or GitLab CI
- Knowledge of monitoring platforms such as Prometheus, Grafana, Datadog, or Splunk
- Strong scripting skills using Python, Bash, or Go
- Excellent troubleshooting, communication, and collaboration skills
Nice to have
- AWS, Azure, or Google Cloud certifications
- Experience with Kubernetes administration
- Knowledge of networking, DNS, and load balancing
- Experience with configuration management tools such as Ansible
- Familiarity with security best practices and compliance standards
Benefits & perks
- Hybrid work flexibility within the United States
- Competitive compensation package
- Medical, dental, and vision insurance
- Paid time off, holidays, and sick leave
- 401(k) retirement savings plan with company match
- Annual performance bonuses
- Professional development and certification reimbursement
- Paid training and technical conference opportunities
- Career growth into Senior Site Reliability Engineer, Platform Engineer, DevOps Architect, Infrastructure Manager, or Cloud Engineering leadership roles
- Collaborative and innovative engineering environment
Why apply
This role is ideal for professionals who enjoy solving complex infrastructure challenges, improving system reliability, and building scalable cloud platforms. You'll have the opportunity to work with modern technologies while advancing your career in cloud engineering, DevOps, and site reliability.
About Careerscape
Careerscape is a staffing and recruiting firm connecting talented professionals with employers across administration, operations, healthcare, finance, technology, customer support, and professional services.