Jobs · Maryland

Site Reliability Engineer

Axle · Frederick, MD · 4 wk ago
Hybrid$140k/yrFull-time

About the role

The Site Reliability Engineer (SRE) role at Axle focuses on modernizing and consolidating a complex multi-cloud environment across AWS, Azure, and GCP. The role involves designing and implementing enterprise-grade monitoring and observability frameworks, establishing and managing Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets to drive reliability improvements. The SRE will also automate workload onboarding and offboarding processes, build proactive detection mechanisms using AIOps and intelligent alerting, and embed ITIL processes into SRE workflows.

Responsibilities

  • Design and implement enterprise-grade monitoring and observability frameworks (metrics, logs, traces) across distributed systems using enterprise Splunk, Grafana, and Open-telemetry tools.
  • Establish and manage SLIs, SLOs, and error budgets to drive reliability improvements.
  • Develop and maintain real-time asset inventory systems across cloud, on-prem, and hybrid environments.
  • Automate workload onboarding and offboarding processes, ensuring standardization and governance.
  • Track system ownership, dependencies, and lifecycle states for operational transparency.
  • Build proactive detection mechanisms using AIOps and intelligent alerting to minimize incident impact.
  • Design and operate scalable, resilient, and secure infrastructure platforms across cloud and hybrid environments.
  • Implement automated compliance tracking and enforcement aligned with organizational and regulatory standards (e.g., NIST, FISMA, FedRAMP).
  • Embed ITIL processes (incident, change, problem, configuration management) into SRE workflows.
  • Build and maintain automated deployment environments and pipelines that enforce security, compliance, and operational standards.
  • Automate provisioning, patching, configuration management, and environment lifecycle.
  • Leverage AI/ML coding assistants and vibe coding practices to rapidly develop automation scripts, tools, and internal platforms.
  • Integrate AI-driven tooling into DevOps pipelines for code quality, security scanning, and operational insights.
  • Lead adoption of AI-enhanced SRE practices, including intelligent remediation and predictive operations.
  • Champion DevOps and SRE practices including Infrastructure as Code, CI/CD, observability, and reliability engineering.
  • Build developer-friendly platforms (“golden paths”) that simplify deployments, reduce friction, and improve velocity.
  • Enable and optimize infrastructure for AI/ML workloads, including data pipelines, storage systems, and inference environments, GPU-enabled and high-performance compute workloads.
  • Build and manage containerized and orchestrated platforms (Docker, Kubernetes).
  • Support cloud migration, modernization, and platform standardization initiatives.
  • Ensure systems meet security, compliance, backup, and disaster recovery requirements.
  • Evangelize and promote best practices in DevOps, SRE, and platform engineering to developer communities.
  • Stay abreast of new technologies in areas such as AIOps, MLOps, cloud computing & deployment, site reliability engineering, infrastructure automation, security best practices, data engineering, etc.

Requirements

  • Total of 6+ years of experience in DevOps/SRE roles with monitoring and observability tools (Prometheus, Grafana, ELK, or cloud-native equivalents) for on-prem and cloud hosted workloads.
  • 4+ years of hands-on Linux experience that includes Ubuntu/CentOS/Red Hat operating systems, containers, dependency management, and administration support.
  • 4+ years of experience automating Infrastructure-as-Code (IaC) deployments to one of the following cloud platforms: Amazon AWS, Google GCP, and Microsoft Azure.
  • 4+ years of experience with CI/CD and automation tools such as Terraform, Ansible, Chef, Puppet, Jenkins, GitHub Actions.
  • Strong scripting skills (Python, Bash, PowerShell or similar).
  • Proficiency in vibe coding and coding assistants to develop scripts, tools, and applications for the DevOps and SRE use cases.
  • Proficiency to debug or troubleshoot and/or deploying SQL and/or NoSQL databases, object storage, web servers, open-source programming stack for Node.JS, R, Python, .NET Core, Java is desired but not mandatory.
  • Willingness to learn new technologies, adopt, and adapt to emerging technologies or needs from a project to a project.
  • Cloud certifications are preferred.
  • Certifications in Grafana, Splunk, Docker, Kubernetes are preferred but optional.

Similar jobs

Site Reliability Engineer

Cracker BarrelTennessee, United States· 1 mo ago
Engineeringapply on cbrlgroup.wd503.myworkdayjobs.com