Site Reliability Engineer III
Apex Systems · Plano, TX · Yesterday
EngineeringContract
Role Overview
We are seeking a Site Reliability Engineer to support the production and operations of critical applications. This role focuses on establishing and improving monitoring to measure end-to-end performance and availability. The ideal candidate will ensure operational readiness by evaluating application performance, reliability, scale, and resiliency, while also identifying and triaging production issues to determine the root cause.
Key Responsibilities
- Perform full-stack triaging of alerts to identify the root cause of application performance and stability issues.
- Work with product owners and other stakeholders to define and track service level objectives (SLOs).
- Design and develop dashboards and reports to communicate key performance metrics.
- Identify opportunities to improve alerting posture and update alerts accordingly.
- Collaborate with the engineering team to understand application architecture and perform single-point-of-failure analysis.
- Create and derive NFR/Workload models to ensure performance and resiliency are considered early in the software development lifecycle.
- Execute performance and chaos tests, analyzing results with APM tools to identify stability issues.
- Document findings, analysis, and results, presenting them to stakeholders.
- Perform analytics on past incidents to understand root causes and implement automation to reduce recurrence.
- Demonstrate proficiency with DevOps tools such as JIRA, BMC Remedy, and ServiceNow.
Required Qualifications
- Experience: 8 to 10 years of information technology experience, with over 6 years on a DevOps, SRE, or performance engineering team.
- Technical Skills:
- Experience triaging production issues using APM tools (e.g., Dynatrace, AppDynamics, Prometheus, Grafana) and log aggregation tools (e.g., Splunk, ELK).
- Strong experience in Java and front-end development (React JS, Angular).
- Proficiency with Apache/Tomcat middleware and Java/RESTful services frameworks.
- Backend database experience with Oracle, SQL Server, or Hadoop.
- Strong scripting skills in Python, UNIX, or Perl/Shell.
- Experience with CI/CD tools such as Bitbucket, JFrog Artifactory, Jenkins, Terraform, and Ansible.
- Knowledge of SRE concepts like SLIs/SLOs and error budgets.
- Experience with the Agile/Scrum methodology.
- Education: A college degree or equivalent work experience is required.
Preferred Qualifications
- Proficiency in system, network, security, and database operations.
- Experience with tools such as Tanium or BMC TrueSight Orchestration.
- Experience with command-line interfaces (CLI), third-party APIs, and integration.
- Server administration experience with Red Hat Enterprise Linux and Windows Server.
- Understanding of developing fault-tolerant solutions and knowledge of horizontal scaling and high availability.
Benefits
- Supplemental benefits including medical, dental, vision, life, disability, and other insurance plans that offer an optional layer of financial protection.
- ESPP (employee stock purchase program) and a 401K program which allows you to contribute typically within 30 days of starting, with a company match after 12 months of tenure.
- HSA (Health Savings Account on the HDHP plan).
- SupportLinc Employee Assistance Program (EAP) with up to 8 free counseling sessions.
- Corporate discount savings program and other discounts.
- On-demand training program, access to certification prep and a library of technical and leadership courses/books/seminars once you have 6+ months of tenure.
- Certification discounts and other perks to associations that include CompTIA and IIBA.
- Dedicated customer service team for Consultants to address questions around benefits and other resources, as well as a certified Career Coach.
- Access to a full list of benefits, programs, support teams and resources within the ‘Welcome Packet’ (available from an Everforth Apex team member).