Senior Reliability Developer
About the role
Responsible for changes and updates to public/private cloud infrastructure. Owns all monitors that support AWN business services and our customers. Implements effective automate using programming languages such as Bash, Python, and JavaScript. Ensures our full infrastructure stack is resilient with day-to-day care and feeding. Writes and manages Terraform modules and configuration trees for deployments into AWS and OpenStack. Manages automated patching solutions for instances deployed across multiple regions to ensure compliance with strict security requirements. Monitors internal and external TLS certificates for expiration, renewals, and re-deployment into their respective environments. Creates and improves service monitoring solutions using Prometheus, Grafana, Zabbix, and various Prometheus exporters (e.g., postgres, node-exporter). Migrates monitoring for legacy production services to a new monitoring system and recreates existing monitoring within Prometheus + Grafana. Collaborates frequently with development teams to implement new features, fixes, or troubleshoot services in lab/production environments. Troubleshoots problems with microservices running within containers, administers containers in Kubernetes clusters, manages cloud components deployed in multiple AWS regions, and maintains CI/CD pipelines. Manages the Cylance AWS Amazon cloud platform through regular administration of Kubernetes cluster upgrades, cluster module upgrades, logging configurations, and resizing of EBS volumes and instances. Troubleshoots issues between interconnected services by working with various teams and resolves issues related to API calls and service unavailability. Creates and maintains detailed documentation for MOPs, service architecture, runbooks, monitoring configurations, and other general resources. Plans and executes service decommissioning in both lab and production environments upon customer or service owner request. Works on an on-call rotation to support business-critical services outside of business hours, responds to escalations, and conducts root cause analysis during incident reviews.
Requirements
- Bachelor’s degree or foreign degree equivalent in Computer Information Systems, or related field and five (5) years of progressive, postbaccalaureate experience in a Technology related role or job offered or related role.
- Experience and/or education must include utilizing Infrastructure-as-Code (IaC) tools including Terraform and configuration management tools including Puppet and Chef to design, implement, and manage multi-region cloud infrastructure deployments across AWS and OpenStack environments.
- Utilizing containerization and orchestration technologies including Docker, Kubernetes, ECS, and EKS to deploy, manage, and troubleshoot microservices architectures in production environments.
- Utilizing Python, Bash, JavaScript, Groovy, and PowerShell scripting languages for automation of infrastructure provisioning, certificate lifecycle management, and automated patching solutions.
- Utilizing monitoring and observability tools including Prometheus, Grafana, Zabbix, AlertManager, and PagerDuty to create dashboards, configure alerts, and implement service health monitoring solutions.
- Utilizing AWS cloud services including EC2, ECS, EKS, ELB, S3, RDS, IAM, Lambda, CloudFormation, VPC, Route53, and CloudWatch to architect, deploy, and manage highly available cloud-based systems.
- Utilizing CI/CD pipeline tools including Jenkins, GitLab, GitHub, Bitbucket, and Gaia to automate testing, deployment processes, and maintain continuous integration workflows.
- Managing Kubernetes cluster operations including cluster upgrades, module upgrades, logging configurations, node scaling, EBS volume management, and pod orchestration across multiple AWS regions.
- Utilizing certificate management and security tools including Keeper for secret management, TLS certificate monitoring, renewal coordination, and automated deployment across production and development environments.
- Troubleshooting distributed systems and microservices communication issues including API failures, service mesh architectures, container networking, inter-service dependencies, and deployment failures.
- Creating and maintaining technical documentation using Atlassian tools including Jira, Confluence, Backstage, and LucidChart for MOPs (Method of Procedures), runbooks, service architecture diagrams, and monitoring configurations.
Skills
- Infrastructure-as-Code (IaC) tools including Terraform and configuration management tools including Puppet and Chef.
- Containerization and orchestration technologies including Docker, Kubernetes, ECS, and EKS.
- Python, Bash, JavaScript, Groovy, and PowerShell scripting languages.
- Monitoring and observability tools including Prometheus, Grafana, Zabbix, AlertManager, and PagerDuty.
- AWS cloud services including EC2, ECS, EKS, ELB, S3, RDS, IAM, Lambda, CloudFormation, VPC, Route53, and CloudWatch.
- CI/CD pipeline tools including Jenkins, GitLab, GitHub, Bitbucket, and Gaia.
- Kubernetes cluster operations including cluster upgrades, module upgrades, logging configurations, node scaling, EBS volume management, and pod orchestration.
- Certificate management and security tools including Keeper for secret management, TLS certificate monitoring, renewal coordination, and automated deployment.
- Troubleshooting distributed systems and microservices communication issues including API failures, service mesh architectures, container networking, inter-service dependencies, and deployment failures.
- Technical documentation creation and maintenance using Atlassian tools including Jira, Confluence, Backstage, and LucidChart.
Benefits
Commensurate with experience.
Pay
Base salary range: $145,018 - $180,000 annually.
Schedule
Flexible.