Jobs · Engineering · Ohio

Site Reliability Engineer (SRE) – II

Huntington National Bank · Columbus, OH · 3 wk ago
EngineeringFull-time

Incident Management

- Handle complex incidents and ensure service uptime. - Lead troubleshooting efforts for high-impact production issues, providing detailed root cause analysis (RCA) and preventative measures. - Participate in on-call rotations, acting as an escalation point for Level 1 SREs during major incidents.

Automation & Infrastructure as Code (IaC)

- Develop and maintain automation scripts and infrastructure using tools like Terraform, Ansible, or CloudFormation. - Implement automation solutions to eliminate manual tasks and improve system reliability, scalability, and performance.

Performance & Scalability

- Analyze system performance and recommend optimizations for scalability and reliability. - Support capacity planning efforts by monitoring system metrics, traffic patterns, and usage trends to predict future resource needs.

System Design & Architecture

- Collaborate with software engineering teams to influence the design of new services and applications, ensuring they are scalable, reliable, and resilient from the start. - Contribute to architectural decisions, ensuring alignment with best practices in fault tolerance, redundancy, and recovery.

Monitoring & Observability

- Build and maintain robust monitoring, alerting, and observability solutions to proactively detect and resolve issues before they impact end users. - Optimize existing monitoring tools (e.g., Prometheus, Grafana, Datadog, Dynatrace) and build custom dashboards for better visibility into system health.

Security & Compliance

- Ensure systems and infrastructure are secure, compliant, and aligned with organizational policies and industry best practices. - Assist with vulnerability management, system patching, and implementing security measures to protect the integrity and availability of services.

Continuous Improvement

- Lead efforts to continuously improve operational processes, tools, and workflows. - Implement and enforce best practices in deployment, monitoring, and incident management to improve overall system reliability and reduce downtime.

Basic Qualification

- Minimum 5 years of experience in site reliability engineering, DevOps, systems administration, or related roles. - Strong experience with Linux/Unix administration and proficiency in scripting (e.g., Python, Bash, Go). - Deep understanding of cloud platforms (AWS, GCP, Azure) and related services (EC2, S3, Lambda, Kubernetes, etc.). - Experience with containerization and orchestration technologies like Docker and Kubernetes. - Proficiency with monitoring and observability tools such as dynatrace, Prometheus, Grafana, Datadog, ELK Stack, or similar platforms. - Strong understanding of networking fundamentals (DNS, HTTP, TCP/IP), load balancing, and CDNs. - Experience with CI/CD tools (Jenkins, GitLab CI, CircleCI) and infrastructure automation (Terraform, Ansible, Puppet). - Familiarity with distributed systems and microservices architecture. - Excellent problem-solving and troubleshooting skills, especially in diagnosing production issues in high-scale environments.

Preferred

- Background in MLOps, data engineering, and/or cloud-native AI deployment. - Strong communication and documentation abilities. - Knowledge of security best practices for AI and cloud infrastructure. - Contributions to open source AI/SRE projects or relevant technical communities. - Proven track record of managing complex infrastructure, troubleshooting production issues, and optimizing system performance.

Similar jobs