Senior SRE Engineer
About the role
This is a 6-month contract role, with a potential to convert to full-time on the Platform Engineering team, working closely with global development, operations, and delivery teams. Platform Engineering owns the reliability, availability, and performance of our infrastructure and services across a multi-account, multi-tenant AWS environment spanning multiple regions.
Responsibilities
Design, implement, and maintain scalable systems for uptime, resilience, and performance, while defining and enforcing SLOs, SLIs, and SLAs with product teams.
Own incident detection, escalation, and resolution, developing playbooks for failure scenarios and leading post-mortem analyses with corrective actions.
Build monitoring, logging, and alerting systems, and develop automation for deployments, maintenance, and toil reduction.
Drive CI/CD improvements and maintain operational documentation, runbooks, and best practices.
Partner with developers on reliability by design, fostering shared responsibility for observability and fault tolerance.
Enforce security policies and standards, collaborating on vulnerability identification and mitigation strategies.
Requirements
AWS Cloud Services: Deep working knowledge across compute, networking, security, storage, and data services in multi-account, multi-region environments.
Kubernetes/EKS: Experience deploying and maintaining Kubernetes with multi-AZ, multi-region, and multi-tenant topologies.
Hybrid Networking: Expertise in VPC design, VPN connectivity, on-prem/hybrid topology, and edge solutions.
Cert/PKI Management: Knowledge of certificate and PKI lifecycle management, including HSM-backed key protection and mTLS.
Security & Cryptography: Strong understanding of encryption, cryptographic protocols, and HSM integration.
Scripting & Automation: Proficiency in Python, Go, and Bash for automation and infrastructure-as-code.
Compliance as Code: Experience using AWS Config to translate policy requirements into automated compliance checks.
Observability Platforms: Hands-on experience with Datadog for metrics, distributed tracing, log management, and APM.
Chaos Engineering: Experience applying chaos engineering practices to proactively validate system resilience.
Preferred Skills & Attributes
Tactical & Operational: Ability to drive team initiatives, own delivery outcomes, and act as a reliability champion.
System Reliability Advocacy: Champion robust system design and operational excellence across engineering and product teams.
Analytical Problem-Solving: Skilled at diagnosing and resolving complex, multi-layer issues in distributed, cloud-native environments.
Cross-Functional Communication: Effective collaborator across development, security, and operations teams in a global organization.
Incident Response Leadership: Experience leading swift resolution of high-severity incidents to minimize downtime.
Culture of Ownership: Fosters accountability and a blameless culture focused on reliability and continuous improvement.
Adaptability: Comfortable problem-solving creatively in fast-moving environments with shifting priorities.