Jobs · Georgia

Site Reliability Engineer

Unum · Dunwoody, GA · Yesterday
$152k–$162k/yrFull-time

About the role

Unum Group seeks Site Reliability Engineers in Atlanta, GA. Applicants who are interested in this position may apply at www.jobpostingtoday.com (Ref #66753) for consideration.

Responsibilities

  • Design, build, and maintain observability, monitoring, and alerting capabilities across consumer and client-facing digital platforms
  • Develop and maintain dashboards that measure availability, latency, error rate, throughput, capacity, MTTR/MBTI, and other reliability metrics
  • Diagnose and troubleshoot distributed systems issues across cloud-based and on-prem services
  • Partner with engineering teams to improve service reliability, reduce operational toil, and mature incident response practices
  • Implement automation for service health checks, performance monitoring, and remediation
  • Manage CI/CD pipeline reliability and deployment quality controls
  • Conduct root-cause analysis and drive long-term corrective actions
  • Collaborate with Run teams to transition monitoring, dashboards, and operational insights into production support processes
  • Provide guidance on service re-platforming, performance improvements, and architectural decisions based on reliability data

Requirements

  • Bachelor’s degree in Computer Science, Engineering, or related field plus 5 years of experience
  • 5 years of experience with the following:
    • Observability and monitoring platforms used to monitor application performance and system health using Dynatrace, AWS CloudWatch, Datadog, Grafana, or Amplitude
    • Working with containerized and cloud-native architectures, including deployment, configuration, and operational support in cloud environments, using AWS
    • Sustaining incident response processes, including participation in on-call rotations, post-incident reviews, and implementation of service-level objectives (SLOs), service-level indicators (SLIs), or service-level agreements (SLAs)
    • Developing scripts or automation to improve system reliability or operational efficiency using Python, Bash, or PowerShell
    • Troubleshooting distributed systems and analyzing performance bottlenecks across multi-tier or microservices-based architectures
    • Collaborating with cross-functional engineering teams, including software engineering, platform, infrastructure, or operations teams, within a DevOps or reliability-focused environment
    • Working with version control systems and collaborative development workflows using GitHub, GitLab, or Bitbucket
  • 4 years of experience with the following:
    • Designing, implementing, or maintaining logging, metrics, and distributed tracing pipelines for enterprise or cloud-based systems
    • Hands-on experience with continuous integration and continuous deployment (CI/CD) tools and pipelines, using GitHub Actions, Jenkins, or Azure DevOps
    • Using infrastructure-as-code or configuration management tools to provision, manage, or maintain environments, including Terraform, AWS CloudFormation, or Ansible

Qualifications

  • Telecommuting w/i worksite. Up to 5% domestic travel
  • 40 hours/week; $152,131 - $162,131 per year

Benefits

  • Healthcare benefits (health, vision, dental)
  • Insurance benefits (short & long-term disability)
  • Performance-based incentive plans
  • Paid time off
  • A 401(k) retirement plan with an employer match up to 5% and an additional 4.5% contribution whether you contribute to the plan or not

Pay

  • $152,131 - $162,131 per year

Schedule

  • 40 hours/week

Similar jobs

Site Reliability Engineer

QualityAIIllinois, United States· Yesterday
Quality Assurance$110k–$120k/yrapply on careers.quality-ai.com