Senior/Staff Site Reliability Engineer - Data Center
PathAI · Boston, MA · 1 mo ago
RemoteRemoteEngineering$166k–$224k/yrFull-time
About the role
We are seeking a senior/staff level Site Reliability Engineer to join our dynamic team. This role will focus on designing, building, and operating our hybrid cloud/on-prem environment.
Responsibilities
- Advancing the state of our operations by implementing SRE best practices - focusing on users, monitoring, and automation.
- Engineering infrastructure patterns for cloud environments in Amazon Web Services - building in security, reliability and scalability.
- Designing, building, and operating our data center to support our rapidly growing Machine Learning team.
- Integrating on-premises datacenter environments with existing cloud infrastructure to create a seamless hybrid cloud environment.
- Improving the reliability and resilience of our infrastructure through root-cause analysis and reviewing gaps in designs, and implementations of our infrastructure.
- Participating in platform on-call rotations and assisting with urgent incident response.
Requirements
- 5+ years of relevant experience.
- Automation: You work hard to eliminate toil by automating everything through scripting, configuration management tools (Ansible), and code (Python/GoLang).
- Moderate experience with monitoring infrastructure with modern observability tools (Datadog/Grafana/Prometheus).
- Experience with infrastructure as code (Terraform/Cloudformation).
- Experience administering physical hardware stacks in production settings (iDRAC/IPMI/Nvidia UFM/Juniper Systems).
- Familiarity with modern network designs and comfort operating across network layers.
- Some experience and opinions on virtualization, containerization, or container orchestration platforms (EKS/ClusterAPI/KVM).
- Operations experience: You’ve managed critical production infrastructure and are familiar with incident response, scaling, and rapid growth related challenges.
Qualifications
- A bachelor's degree in Computer Science or equivalent experience.
- An insatiable intellectual curiosity and the ability to learn quickly in a complex space.
Skills
- Opinionated on storage solutions and how they can be optimized for high performance workloads (Quobyte/S3/FSx/EFS).
- Familiarity with modern network designs and comfort operating across network layers.
- Some experience and opinions on virtualization, containerization, or container orchestration platforms (EKS/ClusterAPI/KVM).
Benefits
- The cash compensation outlined below includes base salary or hourly wage and on-target commission for employees in eligible roles.
- The summary below indicates if an employee in this position is eligible for annual bonus, overtime pay and equity awards.
- Individual compensation packages are tailored based on skills, experience, qualifications, and other job-related factors.
Pay
- Annual Pay Range: $165,750 - $224,450
- Not Overtime Eligible
- Eligible for Equity
Schedule
- Willingness to travel up to 25% of the time.