Senior Platform Engineer
Flexential · Denver Metropolitan Area · 1 wk ago
RemoteRemoteEngineering$150k–$165k/yrFull-time
Key Responsibilities And Essential Job Functions
- Design, develop and operationally manage automated, resilient, high availability, self-healing, secure platforms with native-AI capabilities for IT needs, serving both internal as well as customer business capabilities.
- Develop, and manage the Observability OpenTelemetry Central Backend Stack: Grafana Enterprise, Mimir, Loki, Tempo, and Alertmanager on Kubernetes/RKE2 via Helm and GitLab CI-CD.
- Develop iaC and CI-CD for automated provisiong and deployment, including Terraform modules for Infra/VM/storage provisioning, Ansible AWX playbooks for OS/App bootstrap, ArgoCD and Helm for Kubernetes configuration.
- Develop AIOps capabilities on platforms for e.g Observability use-cases: anomaly detection integrations, event correlation rules in Alertmanager, and synthetic monitoring patterns to reduce alert noise.
- Maintain platform security: Conjur/CyberArk secret injection at runtime, mTLS between stack components, RBAC in Grafana Enterprise.
- Author and maintain Grafana dashboards in JSON/GitLab — facility overview, network health, RED metrics, application telemetry.
Required Qualifications
- DevOps / Automation 5+ years in a production environment.
- Kubernetes (RKE2/k3s).
- Helm chart deployment.
- System services.
- Docker/container.
- LGTM Stack Development and Configuration 4+ years: Grafana, Mimir, Loki, Tempo configuration, tuning, dash-boarding and production operations.
- Prometheus required.
- Senior-level Python / Scripting Frameworks 5+ years.
- Automation scripts.
- Exporter development.
- GitLab pipeline scripting.
- REST API integrations.
- GitOps / CI/CD 5+ years.
- GitLab CI/CD pipeline authoring.
- Terraform and Ansible as primary IaC tools.
- ArgoCD or Flux preferred.
- AIOps / Observability Engineering 2+ years.
- Alertmanager rule authoring.
- Anomaly detection integration.
- Event correlation.
- Noise reduction techniques.
- Working Infrastructure (Linux/VM) Management Knowledge 5+ years.
- Linux administration.
- VMware vCenter/VCF experience.
- Netapp storage management.
- Network fundamentals (SNMP, TCP/IP).
- Secrets Management 2+ years.
- CyberArk/Conjur, HashiCorp Vault, or equivalent.
- Runtime secret injection patterns.
Preferred Qualifications
- Experience and/or knowledge of ITSM processes and workflow automation e.g. Incident & Response Mgmt (IRM), Release mgmt., ServiceNow ITSM integration, alert routing, escalation policy design, SLA-driven on-call workflows.
- Hands-on experience or working knowledge of Boomi integrations PaaS(iPaaS) technologies.
- Experience working with BAS / BMS systems in a Datacenter / OT environment.
- Hands-on experience working with AWS products in a Well-architected Framework and multi-account model to develop various compute, storage, network iaaS and PaaS services for IT applications.