Jobs · Information Technology · New Hampshire

Site Reliability Engineer

TalentAlly · Merrimack, NH · 1 mo ago
Information TechnologyFull-time

About the Role

We are looking for a Site Reliability Engineer who solves operational problems by building software. In this role, you will improve reliability, reduce toil, and enhance production systems by writing code, building automations, and leveraging modern AI-assisted development tools. The position requires a deep understanding of application and infrastructure support and expertise in supporting cloud computing environments. This is a hands-on engineering role, not a traditional support position.

Responsibilities

  • Build automation, scripts, and lightweight tools to eliminate repetitive manual work and improve operational efficiency.
  • Manage Kubernetes cluster administration and troubleshoot Kubernetes-related issues.
  • Participate in on-call rotations, respond to incidents, execute runbooks, and ensure clear communication and handoffs.
  • Continuously analyze system performance in production, troubleshoot consumer-reported issues, and proactively identify areas in need of optimization.
  • Work across teams to lead change in a creative and collaborative manner with business and technical partners.
  • Support 24/7, continuous availability production and managed environments.

Requirements

  • 2+ years of experience in systems and platform operations and technology management.
  • Experience in Cloud computing (Azure and AWS), VMs, Windows, and Linux.
  • Experience with Python scripting and PowerShell skills are highly preferred.
  • Experience with analytics and monitoring tools such as Grafana, Splunk, and Datadog.
  • Experience with continuous integration tools, such as Jenkins and AWX.
  • Good understanding of software architecture to empower software developers and engineers to build platforms with greater resiliency and fault tolerance.
  • Knowledge of SRE (Site Reliability Engineering) principles: resiliency, observability, and governance gating.
  • Good to have knowledge of networking, firewalls, and load balancers.
  • Knowledge of best practices for IT operations in an always-on, always-available service model.
  • Great communication, collaboration, and interpersonal skills.

Optional Qualifications

  • AWS or Azure-related certifications.

About the Team

As a member of the Fidelity Health Recordkeeping Production Support and WI CloudOps team, you will collaborate with other technology professionals to support WI Technology - Cloud Platform solutions. The team brings together individuals from diverse technical backgrounds, and the role offers exposure to a wide range of challenges. Ideal candidates will have a background in Site Reliability with a strong desire to expand into other domains, or prior experience as a Site Reliability Engineer. We are looking for a systems-thinking Site Reliability Engineer who has helped teams scale through production insights, operational automation, developer enablement, real-time metrics, and continuous improvement.

Schedule

Fidelity is transitioning to a full-time onsite working model through a phased rollout across regions and roles. Currently, some roles and locations require 100% onsite presence, while others require less. Onsite expectations are likely to evolve as the rollout continues. This transition does not apply to fully remote roles.

Similar jobs

Site Reliability Engineer

Cracker BarrelTennessee, United States· 1 mo ago
Engineeringapply on cbrlgroup.wd503.myworkdayjobs.com