Jobs · Information Technology · Connecticut

Reliability Engineer

The Hartford · Hartford, CT · 1 mo ago
On-siteInformation Technology$91k–$137k/yrFull-time

Responsibilities

  • Aid in the use of best-in-class software engineering standards and design practices for instrumenting code/application technology stack to enable the generation of relevant metrics on overall technology health.
  • Support the architecture and software engineering teams in influencing the technical strategy for the organization, considering its cross-functional impacts, integration across the organization, and architecture rationalization.
  • Lead a technical team for applications supported, demonstrating depth and breadth of knowledge in technologies, applications, integration, interfaces, and business domains.
  • Develop effective tooling, alerts, and response mechanisms to identify and address reliability risks using automation to support problem prevention, detection, mitigation, and resolution.
  • Enhance the delivery flow by engineering the appropriate solutions to increase delivery speed while adhering to technology standards for sustained reliability.
  • Implement preventative controls and drive increased automation and self-healing capabilities.
  • Promote and implement innovative solutions.
  • Ensure operational excellence by collaborating to drive the triaging and service restoration of all high impact incidents to minimize mean time to service restoration and impact to the business.
  • Show end-to-end ownership by partnering with infrastructure teams to design and implement intelligent incident routing, enhanced monitoring/alerting capabilities, and automated service restoration processes.
  • Take proactive measures to prevent high impactful incidents and maintain the continuity of Hartford and third-party assets supporting a business function.
  • Keep IT application and infrastructure metadata repositories current.
  • Research and implement AI-based anomaly detection to predict infrastructure failures and automate preventive measures.
  • Develop AI-powered troubleshooting copilots and LLM-driven operational assistants to accelerate incident resolution and root cause analysis.
  • Implement AI/ML-based runbooks to automate system recovery and optimize operational efficiency.

Qualifications

  • Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field.
  • 3+ years of experience in Infrastructure Engineering, Site Reliability Engineering (SRE), or DevOps.
  • Hands-on experience with observability tools: Splunk, Dynatrace, CloudWatch.
  • Deep knowledge of Infrastructure as Code (IaC) with Terraform, CloudFormation.
  • Proven ability to optimize CI/CD pipelines, automate deployments, and enforce DevSecOps best practices.
  • Expertise in cloud platforms (AWS) and Kubernetes-based microservices environments.
  • Strong proficiency in Python, Java for infrastructure automation and tooling development.
  • Experience in AI/ML frameworks for observability, predictive failure detection, and AI-driven troubleshooting desirable.
  • Knowledge of open-source database technologies is beneficial.
  • Demonstrated experience working within Agile frameworks and methodologies.
  • Excellent analytical, problem-solving, and interpersonal skills.

Similar jobs

Reliability Engineer

Decision Technologies, Inc.Arlington, VA· 3 wk ago
Information Technologyapply on ats.rippling.com

Reliability Engineer

Pratt & WhitneyAsheville, NC· 1 mo ago
Information Technologyapply on careers.rtx.com

Reliability Engineer

SyngentaOmaha, NE· 2 mo ago
Supply Chain$250k/yrapply on jobs.smartrecruiters.com

Reliability Engineer

I-care Reliability Inc.Pennsylvania, United States· 1 mo ago
Engineeringapply on icare-usa.cvw.io

Reliability Engineer

James HardieUnited States· 1 mo ago
RemoteEngineering$125k–$145k/yrapply on careers.jameshardie.com