Site Reliability Engineer II
About the role
The Team You Will Join: As Our Site Reliability Engineering Function Within The Human Risk Platform Pillar, a Team That Sits At The Center Of How Mimecast Builds, Ships, And Operates Its Cloud Platform.
The Role: As a Site Reliability Engineer, you will be a hands-on engineer who stands up infrastructure, debugs it when it misbehaves, and turns both activities into repeatable, self-service patterns for the engineering teams we support.
Responsibilities
- Stand up and evolve AWS infrastructure using Terraform, with a strong bias toward reusable modules and paved-road patterns over bespoke solutions.
- Practical IAM least privilege, secrets/certs handling and testability a must.
- Operate and improve Kubernetes-based workloads deployments, scaling, networking, and the platform glue that makes them boring to run.
- Capacity signals and basic performance triage.
- Build and maintain CI/CD pipelines (GitHub Actions preferred) that give engineering teams fast, safe, auditable paths to production.
- Progressive delivery (canary/blue-green), rollback discipline.
- Partner with product squads to enable self-service access to their own infrastructure, databases, and pipelines with the guardrails, auditability, and standards that make that access safe.
- Debug production issues across the stack: networking, DNS, certificates, container orchestration, CI pipelines, and application-level behavior.
- Instrument services with appropriate logs, metrics, and traces, and help teams adopt the observability standards our platform teams define.
- Contribute to runbooks, automation, and standards that reduce toil and one-off work — if you fix it twice, codify it the third time.
- Participate in on-call; lead or co-lead incidents and write blameless postmortems with concrete follow up.
- Use AI tooling pragmatically and at a professional level: to accelerate code generation, infrastructure design, debugging, documentation, and review.
Requirements
- Hands-on Kubernetes experience: You can deploy, operate, and debug workloads on Kubernetes.
- Terraform experience: You have written and maintained non-trivial Terraform.
- Familiarity with setting up and maintaining CI and deployment automation. Experience with GitHub Actions is strongly preferred; experience with Jenkins, GitLab CI, or similar is transferable.
- A working understanding of observability fundamentals, logs, metrics, traces and how they are used during incident response.
- Networking knowledge: Security groups, TLS certificates, DNS, load balancing, and how traffic flows through a cloud environment.
- Experience with AWS or similar cloud platforms.
- Experience working with relational databases is an asset, PostgreSQL is ideal but not required.
- Experience in software development as a full stack developer is a strong asset.
Skills
- AI-First Engineering Role: Mimecast is an AI-first engineering organization. We expect every engineer to use modern AI tools (coding assistants, reasoning models, agentic tools) as a core part of their daily workflow for design, implementation, debugging, review, and documentation.
Benefits
Ways of Working: You collaborate well across teams and can work with product engineers, platform teams, and security peers without friction. You have a bias for action and problem-solving, and you prefer shipping small, frequent improvements over big-bang changes. You are comfortable saying “let’s standardize this” rather than building a one-off, even when the one-off would be faster today. You communicate clearly with the team in both spoken and written form.
Pay
The base salary range for this position is $124,000−$186,000 plus benefits. This range represents the minimum and maximum new hire compensation for this role. The position may also be eligible for incentive plans and additional benefits, in accordance with company policy and local regulations.
Schedule
We provide you with the flexibility to live balanced, healthy lives through our hybrid working model that champions both collaborative teamwork and individual flexibility. Employees are expected to come to the office at least two days per week, because working together in person: Fosters a culture of collaboration, communication, performance and learning. Drives innovation and creativity within and between teams. Introduces employees to priorities outside of their immediate realm. Ensures important interpersonal relationships and connections with one another and our community!