Kubernetes Engineer / Hybrid / Chelmsford, MA
This company is a leading provider of access control and security management solutions, delivering innovative technology that helps organizations protect people, assets, and facilities around the world. As the company continues its cloud-native transformation, they are investing heavily in modern DevOps practices, Kubernetes, CI/CD automation, observability, and infrastructure modernization.
About the role
This is an opportunity to work across Azure, Digital Ocean, Kubernetes, TeamCity, Docker, Prometheus, Grafana, and a broad cloud-native ecosystem while helping shape the future of the platform. This is not a maintenance-focused DevOps position. We are looking for a technology enthusiast who enjoys researching new tools, automating everything possible, and driving continuous improvement. You'll have ownership over engineering environments, deployment tooling, CI/CD pipelines, and observability while collaborating closely with teams supporting production AKS environments. If you enjoy solving complex infrastructure challenges, working directly with customers on real-world deployments, and influencing technical direction, this role offers significant visibility, autonomy, and growth opportunities.
Responsibilities
- Own engineering environments across Azure and Digital Ocean, including development, QA, staging, and performance testing environments
- Design and automate infrastructure provisioning to support on-demand testing and validation environments
- Build, maintain, and improve customer-facing deployment tooling and on-premises installation processes
- Support both Kubernetes and legacy Docker Swarm deployment models for customer environments
- Own and enhance TeamCity CI/CD pipelines, release automation, testing integrations, and security scanning workflows
- Partner with the Hosting Team supporting production AKS infrastructure and assist with troubleshooting and root cause analysis
- Build and maintain observability solutions using Prometheus, Grafana, structured logging, and alerting tools
- Execute performance and load testing initiatives to improve platform scalability and reliability
- Strengthen infrastructure security and contribute to compliance initiatives, including SOC 2-related activities
- Develop runbooks, installation guides, incident response documentation, and engineering knowledge base materials
- Participate in production on-call rotations and incident response efforts
- Research, evaluate, and implement emerging DevOps tools and cloud-native technologies that improve engineering efficiency and platform reliability
Tech Breakdown
- 40% Linux Administration & Infrastructure Engineering
- 25% Kubernetes, Docker & Container Platform Management
- 15% CI/CD Pipelines & Release Automation (TeamCity)
- 10% Observability, Monitoring & Performance Testing
- 10% Customer Deployment Support & Technical Collaboration
Daily Responsibilities
- 80% Hands On
- 5% Management Duties
- 15% Team Collaboration
Requirements
Required Skills & Experience
- Deep Linux expertise, including system administration, troubleshooting, performance tuning, and containerized workloads
- Strong Docker experience, including knowledge of Docker Swarm environments
- Production Kubernetes administration and troubleshooting experience
- Experience supporting on-premises software deployments and customer-hosted infrastructure
- CI/CD pipeline experience with TeamCity, GitHub Actions, GitLab CI, Jenkins, or similar tools
- Experience building deployment packages, installers, or deployment automation for customer environments
- Hands-on experience with monitoring, observability, logging, and alerting platforms
- Experience with load testing and performance testing tools such as k6, Locust, or JMeter
- Understanding of messaging systems such as RabbitMQ, NATS, or similar technologies
- Relational database experience, preferably PostgreSQL
- Strong scripting and automation experience with Python, Bash, Go, or similar languages
- Experience with Git-based source control workflows
- Solid networking knowledge including DNS, TLS, VPNs, load balancing, and firewall concepts
- Strong technical documentation and communication skills
- Customer-facing experience supporting installations, upgrades, or technical deployments
Desired Skills & Experience
- Prometheus and Grafana implementation and administration experience
- Infrastructure-as-Code experience using Terraform, Ansible, or similar tools
- Experience with Mend.io or comparable SAST/SCA security platforms
- Redis and Elasticsearch administration experience
- Azure cloud experience and Azure certifications (AZ-104, AZ-400)
- Digital Ocean infrastructure experience
- Experience supporting air-gapped or highly restricted customer environments
- Familiarity with SOC 2 compliance and security best practices
- Experience with AI-assisted engineering tools such as GitHub Copilot or Claude
- Exposure to physical security, access control, or IoT-related technology platforms
- Experience introducing and driving adoption of modern DevOps tooling and best practices