Site Reliability Engineer II
KFC · Kansas, United States · 1 wk ago
HybridFull-time
About the role
The SRE II leads SRE initiatives, including defining and meeting SLI/SLOs for critical applications, deploying monitoring and logging tools, and creating automation to reduce toil. This role drives blameless RCAs to improve KFC’s systems and collaborates cross-functionally with product owners to establish SLI/SLOs. The SRE II also designs and builds cloud-based infrastructure observability to enhance performance, cost, and reliability, and serves as the team’s expert for IaC in containerized environments.
Responsibilities
- Lead the SRE team’s goal of eliminating toil through automation of processes and tool creation.
- Collaborate with Platform Engineering, Development, and Product teams to design, deploy, and maintain cloud-native applications and infrastructure using Infrastructure as Code (IaC).
- Monitor KFC environments proactively using tools like DataDog to identify and address issues in production.
- Participate in a rotational on-call system to support infrastructure 24/7/365.
- Provide SRE support for KFC platforms through troubleshooting, analysis, and technical resolution of complex distributed systems.
- Deliver root cause analysis and corrective actions after critical incidents within defined SLAs.
- Collaborate with Product Owners to determine SLI/SLOs for projects and design dashboards to track these standards.
- Configure monitoring and logging tools to generate alerts about system and application health.
Requirements
- 5+ years of IT experience.
- 1+ years in a professional technical role with multi-cloud experience (Azure, AWS, GCP preferred).
- 1+ years of experience with bash and/or Linux command-line utilities.
- 1+ years of experience in Python, Go, REST APIs, or GraphQL.
- 2+ years of experience with observability tools such as DataDog, Elastic, or Dynatrace.
- Expertise in CI/CD best practices and methodologies (e.g., GitLab).
- Experience with DevOps-based IaC toolkits like Terraform or Ansible.
- Working experience with incident and problem tracking systems (e.g., ServiceNow, Jira).
- Excellent collaboration skills across engineering functions, business leaders, and vendors.
- Strong verbal and written communication, troubleshooting, and analytical skills.
- Understanding of large-scale infrastructure environments and current industry trends.
- Proven ability to work autonomously with a focus on long-term results.
Bachelor’s Degree preferred.
Pay
Salary Range: $95,700 - $120,000
Benefits
- Medical, dental, vision, legal, and accidental death and dismemberment insurance.
- Flexible Spending Account (FSA) and Health Savings Account (HSA) options (depending on enrolled medical plan).
- Short-term disability, long-term disability, and life insurance.
- 401(k) plan enrollment.
- 4 weeks of vacation, paid sick leave, and 10 paid holidays.
- Floating day off and half-day Fridays year-round.
- 2 paid days for volunteer time each calendar year.