Site Reliability Engineer
Unum · Dunwoody, GA · Yesterday
$152k–$162k/yrFull-time
About the role
Unum Group seeks Site Reliability Engineers in Atlanta, GA. Applicants who are interested in this position may apply at www.jobpostingtoday.com (Ref #66753) for consideration.
Responsibilities
- Design, build, and maintain observability, monitoring, and alerting capabilities across consumer and client-facing digital platforms
- Develop and maintain dashboards that measure availability, latency, error rate, throughput, capacity, MTTR/MBTI, and other reliability metrics
- Diagnose and troubleshoot distributed systems issues across cloud-based and on-prem services
- Partner with engineering teams to improve service reliability, reduce operational toil, and mature incident response practices
- Implement automation for service health checks, performance monitoring, and remediation
- Manage CI/CD pipeline reliability and deployment quality controls
- Conduct root-cause analysis and drive long-term corrective actions
- Collaborate with Run teams to transition monitoring, dashboards, and operational insights into production support processes
- Provide guidance on service re-platforming, performance improvements, and architectural decisions based on reliability data
Requirements
- Bachelor’s degree in Computer Science, Engineering, or related field plus 5 years of experience
- 5 years of experience with the following:
- Observability and monitoring platforms used to monitor application performance and system health using Dynatrace, AWS CloudWatch, Datadog, Grafana, or Amplitude
- Working with containerized and cloud-native architectures, including deployment, configuration, and operational support in cloud environments, using AWS
- Sustaining incident response processes, including participation in on-call rotations, post-incident reviews, and implementation of service-level objectives (SLOs), service-level indicators (SLIs), or service-level agreements (SLAs)
- Developing scripts or automation to improve system reliability or operational efficiency using Python, Bash, or PowerShell
- Troubleshooting distributed systems and analyzing performance bottlenecks across multi-tier or microservices-based architectures
- Collaborating with cross-functional engineering teams, including software engineering, platform, infrastructure, or operations teams, within a DevOps or reliability-focused environment
- Working with version control systems and collaborative development workflows using GitHub, GitLab, or Bitbucket
- 4 years of experience with the following:
- Designing, implementing, or maintaining logging, metrics, and distributed tracing pipelines for enterprise or cloud-based systems
- Hands-on experience with continuous integration and continuous deployment (CI/CD) tools and pipelines, using GitHub Actions, Jenkins, or Azure DevOps
- Using infrastructure-as-code or configuration management tools to provision, manage, or maintain environments, including Terraform, AWS CloudFormation, or Ansible
Qualifications
- Telecommuting w/i worksite. Up to 5% domestic travel
- 40 hours/week; $152,131 - $162,131 per year
Benefits
- Healthcare benefits (health, vision, dental)
- Insurance benefits (short & long-term disability)
- Performance-based incentive plans
- Paid time off
- A 401(k) retirement plan with an employer match up to 5% and an additional 4.5% contribution whether you contribute to the plan or not
Pay
- $152,131 - $162,131 per year
Schedule
- 40 hours/week