Senior DevOps Engineer
RBA, Inc. · Minnesota, United States · Yesterday
RemoteRemoteConsultingContract
What You’ll Do
- Implement and roll out core platform capabilities (Kubernetes-based runtime, build/deploy tooling, observability) across teams and environments.
- Create reusable templates, Helm charts, pipeline definitions, and IaC modules that teams can adopt with minimal friction.
- Onboard scrum teams onto the platform - migrating workloads, standardizing configuration, and documenting self-service paths.
- Drive consistent adoption of platform standards while accommodating legitimate team-specific needs.
- Own day-to-day health, configuration, and lifecycle of the platform and its supporting infrastructure.
- Plan and execute infrastructure and platform upgrades (Kubernetes versions, node pools, runtimes, agents, tooling) with minimal disruption.
- Manage platform configuration changes through controlled, auditable processes.
- Provide responsive support to engineering teams using the platform, acting as the escalation point for platform-related issues.
- Manage SSL/TLS certificate renewals and rotation, secret/credential rotation, patching, and capacity adjustments on a recurring, reliable cadence.
- Plan and carry out infrastructure upgrades and maintenance windows, including rollback planning and stakeholder communication.
- Maintain and tune platform configuration for scalability, resiliency, performance, and cost efficiency.
- Lead incident response for production issues - triage, mitigation, coordination, and resolution - and run blameless post-incident reviews with concrete follow-up actions.
- Improve observability, monitoring, and alerting so issues are caught proactively rather than reported by users.
- Design, implement, and maintain automated CI/CD pipelines covering build, test, release, and deployment across Windows and Linux workloads.
- Enable scrum teams to own their pipelines through shared templates, reusable steps, and clear documentation.
- Promote sound source control, branching, and release management practices.
- Embed security and compliance controls directly into pipelines (DevSecOps) - scanning, policy gates, and secrets handling.
- Continuously reduce lead time, deployment friction, and manual steps in the delivery lifecycle.
- Provide Level-3 support for complex production and platform incidents.
- Identify and act on opportunities to improve system stability, operational maturity, and self-service.
- Optimize platforms and processes based on measurable outcomes (deployment frequency, change failure rate, MTTR, cost).
- Maintain strong, current technical documentation and runbooks.
- Partner with Application Architects and engineering teams to align platform and DevOps solutions with product needs.
- Mentor and coach engineers on DevOps practices, pipeline ownership, and operational discipline.
- Promote a culture of automation, ownership, and continuous improvement.
- Research emerging tools and practices that improve the delivery lifecycle, and introduce cost-effective solutions that increase speed, quality, and reliability.
- Active AI practitioner- leverages AI-assisted tooling (Claude, Cursor, Copilot) to accelerate engineering and operational work.
What You Bring
- At least 5 years of hands-on DevOps, platform, or infrastructure engineering experience within distributed systems or large enterprises.
- Bachelor’s degree in Computer Science or related discipline, or equivalent work experience.
- Kubernetes administration and containerization (Docker, Helm), including platform upgrades and lifecycle management.
- CI/CD orchestration across Windows and Linux environments (TeamCity, Octopus).
- Infrastructure automation and scripting (Terraform, Ansible, PowerShell, Bash).
- Cloud infrastructure administration (Azure / AWS).
- Log collection and dashboarding (ELK, New Relic).
- Demonstrated experience operating production systems and leading incident response.
- Nice to have:
- Experience rolling out internal developer platforms or self-service tooling across multiple teams.
- Familiarity with policy-as-code, secrets management, and DevSecOps tooling.
- Experience defining and tracking delivery/reliability metrics (DORA or similar).