Jobs · Engineering

3114 - Sr. Site Reliability Engineer

Intertech, Inc. · Minnesota, United States · 1 wk ago
EngineeringContract

Remote long-term contract. Must be a US Citizen for this client.

About the role

The Sr Site Reliability Engineer will architect, develop, and maintain cloud environments in both the commercial and government cloud. The role will work closely with software engineers, architects, and DevOps engineers to architect and maintain a secure, resilient, and high-performance cloud infrastructure.

Responsibilities

  • Build, maintain, and operate IaaS and PaaS infrastructure in Azure commercial and government clouds.
  • Work closely with development teams to identify and measure SLOs, SLAs, and SLIs.
  • Act as a strong contributor to the development of platform services, including architecture, provisioning, configuration, deployment, and support.
  • Perform integrations with central logging, metrics dashboards, instrumentation, incident monitoring, and management.
  • Build, integrate, and administer systems and tools that enable engineering teams to observe their applications in production with autonomy (Dashboards, APMs).
  • Support software and/or cloud infrastructure in an on-call rotation.
  • Assist with identification and remediation of technical problems at the root cause by continuously implementing automation, self-healing, and real-time monitoring for production systems.
  • Maintain and improve operational tooling, frameworks, and build frameworks that test the performance and resiliency of platform services/tools.
  • Automate alerts for metrics on performance, cost, vulnerabilities, risk, and compliance violations.
  • Improve processes and champion automation of any manual items around support.

Requirements

  • 4+ years of experience working within an SRE engineer/cloud platform role.
  • Experience leveraging AI tools in the software development (or product) lifecycle to improve quality and efficiency.
  • Expert knowledge of a cloud service provider.
  • Expert knowledge and hands-on production experience in Kubernetes (bare metal or managed) cluster setup and management.
  • Experience with infrastructure as code (IaC) tools like Terraform, Pulumi.
  • Experience with Kubernetes deployment tools like Helm, ArgoCD, Flux.
  • Experience with monitoring tools (Dynatrace, Azure Monitor, Splunk, Grafana, Prometheus).
  • Experience with the Elastic Stack or ELK stack (Elasticsearch, Logstash, and Kibana).
  • Strong awareness of networking and internet protocols.
  • Understanding of identity and access management (IAM).
  • Experience supporting infrastructure in production cloud environments.
  • Knowledge of encryption, Public Key Infrastructure (PKI), and understanding of OWASP.
  • Experience working with RESTful services.
  • Familiarity with IDEs and source control tools like Visual Studio Code and Git.

Preferences

  • Bachelor’s Degree in Computer Science, Information Technology, Software Engineering, Math, Physics.
  • Master’s Degree with coursework focused on advanced algorithms, mathematics in computing, data structures, or related field.
  • Expert knowledge of Azure.
  • Demonstrated passion for infrastructure automation.
  • Ability to prioritize work in a fast-paced environment.

Similar jobs