Early Career - Site Reliability Engineer
Description
As a Retail Site Reliability Engineer (SRE) within Client's Site Reliability Center, you will combine your software and systems expertise to manage applications and create innovative, automated solutions to simplify operations, eliminate toil, and increase the reliability and availability of our critical applications and business services.
Objective
Run the production environment by monitoring availability and taking a holistic view of system health
Build software and systems to manage platform infrastructure and applications
Improve reliability, quality, and time-to-market of our suite of software solutions
Measure and optimize system performance, with an eye toward pushing our capabilities forward, getting ahead of customer needs, and innovating for continual improvement
Responsibilities
- Gather and analyze metrics from operating systems as well as applications to assist in performance tuning and fault finding
Partner with development teams to improve services through rigorous testing and release procedures
Create sustainable systems and services through automation and uplifts
Balance feature development speed and reliability with well-defined service-level objectives
Partner with technology teams across the enterprise to establish SRE best practices and automated solutions with a focus on operational excellence
Identify opportunities to evangelize adoption for greater self-healing and resiliency patterns
Troubleshoot priority incidents and participate in blameless post-mortems
Perform analytics on previous incidents and usage patterns to better predict issues and take proactive actions
Participate in 24x7 on-call rotations and escalation workflows
Requirements
- Experience and detailed knowledge in the following Engineering and support of micro app/service architectures
Working knowledge of web services technologies such as SOAP, JSON and REST
Application Platforms such as OpenShift, Docker and Kubernetes
Development Frameworks such as React, Node.js and Spring Boot
Expertise in database technologies including SQL Server, MySQL, Oracle and Mongo
Understanding of distributed tracing and monitoring tools such as Dynatrace, Jaeger and Humio
Experience with Kafka event streaming
Strong working knowledge of modern development technologies and tools such as Agile, CI/CD, Git and Jenkins
Strong working knowledge of Internet protocols such as HTTP, TCP/UDP