Site Reliability Engineer II
Apex Systems · Chandler, AZ · 1 mo ago
EngineeringContract
About the role
The Site Reliability Engineer (SRE) will manage the entire lifecycle of cloud services, ensuring high levels of service reliability through monitoring, troubleshooting, and automation.
Responsibilities
- Maintain live services by measuring and monitoring availability, latency, and overall system health.
- Troubleshoot issues across the entire technology stack, including hardware, software, applications, and networks.
- Perform deep-dive analysis into both systemic and latent reliability issues, partnering with engineering and operations teams to implement fixes.
- Drive standardization efforts across multiple disciplines and services in conjunction with other SREs throughout the organization.
- Identify and pursue opportunities to improve automation for cloud services.
- Scope and create automation for the deployment, management, and visibility of services.
Requirements
- Experience: 5+ years of experience working with Unix/Linux Server platforms.
- Technical Skills: Proficiency in Terraform, Shell scripting, Java, Python, and Ansible development.
- Experience with the full lifecycle of cloud services.
Preferred Qualifications
- Ability to troubleshoot complex issues across hardware, software, application, and network layers.
- Experience performing deep dives into systemic and latent reliability issues and collaborating with cross-functional teams to implement fixes.
- A track record of identifying and driving opportunities to improve automation for cloud infrastructure and services.
Benefits
Everforth Apex offers a comprehensive benefits package including medical, dental, vision, life, disability, and other insurance plans, an ESPP, a 401K program, an HSA, a SupportLinc Employee Assistance Program (EAP), a corporate discount savings program, and other discounts. Professional development includes an on-demand training program, certification prep, and access to technical and leadership courses/books/seminars.