Reliability Engineer
The Hartford · Hartford, CT · 1 mo ago
On-siteInformation Technology$91k–$137k/yrFull-time
Responsibilities
- Aid in the use of best-in-class software engineering standards and design practices for instrumenting code/application technology stack to enable the generation of relevant metrics on overall technology health.
- Support the architecture and software engineering teams in influencing the technical strategy for the organization, considering its cross-functional impacts, integration across the organization, and architecture rationalization.
- Lead a technical team for applications supported, demonstrating depth and breadth of knowledge in technologies, applications, integration, interfaces, and business domains.
- Develop effective tooling, alerts, and response mechanisms to identify and address reliability risks using automation to support problem prevention, detection, mitigation, and resolution.
- Enhance the delivery flow by engineering the appropriate solutions to increase delivery speed while adhering to technology standards for sustained reliability.
- Implement preventative controls and drive increased automation and self-healing capabilities.
- Promote and implement innovative solutions.
- Ensure operational excellence by collaborating to drive the triaging and service restoration of all high impact incidents to minimize mean time to service restoration and impact to the business.
- Show end-to-end ownership by partnering with infrastructure teams to design and implement intelligent incident routing, enhanced monitoring/alerting capabilities, and automated service restoration processes.
- Take proactive measures to prevent high impactful incidents and maintain the continuity of Hartford and third-party assets supporting a business function.
- Keep IT application and infrastructure metadata repositories current.
- Research and implement AI-based anomaly detection to predict infrastructure failures and automate preventive measures.
- Develop AI-powered troubleshooting copilots and LLM-driven operational assistants to accelerate incident resolution and root cause analysis.
- Implement AI/ML-based runbooks to automate system recovery and optimize operational efficiency.
Qualifications
- Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field.
- 3+ years of experience in Infrastructure Engineering, Site Reliability Engineering (SRE), or DevOps.
- Hands-on experience with observability tools: Splunk, Dynatrace, CloudWatch.
- Deep knowledge of Infrastructure as Code (IaC) with Terraform, CloudFormation.
- Proven ability to optimize CI/CD pipelines, automate deployments, and enforce DevSecOps best practices.
- Expertise in cloud platforms (AWS) and Kubernetes-based microservices environments.
- Strong proficiency in Python, Java for infrastructure automation and tooling development.
- Experience in AI/ML frameworks for observability, predictive failure detection, and AI-driven troubleshooting desirable.
- Knowledge of open-source database technologies is beneficial.
- Demonstrated experience working within Agile frameworks and methodologies.
- Excellent analytical, problem-solving, and interpersonal skills.