Jobs · Massachusetts

Senior SRE Engineer

Akoya · Boston, MA · 4 days ago
Hybrid$55–$75/hrContract

About the role

This is a 6-month contract role, with a potential to convert to full-time on the Platform Engineering team, working closely with global development, operations, and delivery teams. Platform Engineering owns the reliability, availability, and performance of our infrastructure and services across a multi-account, multi-tenant AWS environment spanning multiple regions.

Responsibilities

  • Design, implement, and maintain scalable systems for uptime, resilience, and performance, while defining and enforcing SLOs, SLIs, and SLAs with product teams.

  • Own incident detection, escalation, and resolution, developing playbooks for failure scenarios and leading post-mortem analyses with corrective actions.

  • Build monitoring, logging, and alerting systems, and develop automation for deployments, maintenance, and toil reduction.

  • Drive CI/CD improvements and maintain operational documentation, runbooks, and best practices.

  • Partner with developers on reliability by design, fostering shared responsibility for observability and fault tolerance.

  • Enforce security policies and standards, collaborating on vulnerability identification and mitigation strategies.

Requirements

  • AWS Cloud Services: Deep working knowledge across compute, networking, security, storage, and data services in multi-account, multi-region environments.

  • Kubernetes/EKS: Experience deploying and maintaining Kubernetes with multi-AZ, multi-region, and multi-tenant topologies.

  • Hybrid Networking: Expertise in VPC design, VPN connectivity, on-prem/hybrid topology, and edge solutions.

  • Cert/PKI Management: Knowledge of certificate and PKI lifecycle management, including HSM-backed key protection and mTLS.

  • Security & Cryptography: Strong understanding of encryption, cryptographic protocols, and HSM integration.

  • Scripting & Automation: Proficiency in Python, Go, and Bash for automation and infrastructure-as-code.

  • Compliance as Code: Experience using AWS Config to translate policy requirements into automated compliance checks.

  • Observability Platforms: Hands-on experience with Datadog for metrics, distributed tracing, log management, and APM.

  • Chaos Engineering: Experience applying chaos engineering practices to proactively validate system resilience.

Preferred Skills & Attributes

  • Tactical & Operational: Ability to drive team initiatives, own delivery outcomes, and act as a reliability champion.

  • System Reliability Advocacy: Champion robust system design and operational excellence across engineering and product teams.

  • Analytical Problem-Solving: Skilled at diagnosing and resolving complex, multi-layer issues in distributed, cloud-native environments.

  • Cross-Functional Communication: Effective collaborator across development, security, and operations teams in a global organization.

  • Incident Response Leadership: Experience leading swift resolution of high-severity incidents to minimize downtime.

  • Culture of Ownership: Fosters accountability and a blameless culture focused on reliability and continuous improvement.

  • Adaptability: Comfortable problem-solving creatively in fast-moving environments with shifting priorities.

Similar jobs