Jobs · Engineering

Principal Site Reliability Engineer

Expel · United States · 2 wk ago
RemoteRemoteEngineering$195k–$225k/yrInternship

About the role

Your passion for uptime was forged from experience in production and refined through incident response. As an Expel Principal Site Reliability Engineer, you are a protector, champion, and leader of Expel's reputation for service reliability. You understand that operational reliability is a shared mission across all of engineering, making it as easy as possible for Expel to achieve that mission. You collaborate with architects and product stakeholders, mentor junior SREs, and apply your dedication to reliability within the broader SRE community to maintain outstanding reliability standards in the cloud native ecosystem.

What Expel Can Do For You

  • Provide an opportunity to grow and maintain reliability-focused platform features within a cloud native engineering platform using modern infrastructure and tooling (Kubernetes/GKE/EKS runtime, Hashicorp toolset, etc).
  • Offer a mission you can get behind: stopping evil hackers so our customers can focus on their business.
  • Include you in a company focused on creating opportunities to do interesting work and creating space for employees to learn and grow.
  • Provide an opportunity to contribute to a best-in-class product.
  • Offer a leadership team that’s embraced modern Site Reliability principles, as outlined by Google and other industry leaders.

Responsibilities

  • Lead project work to build and maintain platform features that cut across the Expel product’s reliability, networking, and cloud infrastructure.
  • Contribute by pushing IaC commits daily, with occasional opportunities to write and test application code in Python, Golang, and JavaScript.
  • Mentor and motivate service owners on how to use the platform to deploy, measure, monitor, and operate their own services at scale.
  • Participate in a weekly support rotation that includes taking the on-call pager and providing nearly on-demand working-hours support to platform users.
  • Lead incident response, triage, and root cause analysis support.

Requirements

  • A passion for learning and improving your work product.
  • Significant experience operating Kubernetes within highly distributed environments.
  • Experience running systems in GCP or AWS.
  • Exposure to monitoring and observability infrastructure and standard methodologies.
  • An understanding of infrastructure-as-code practices, tools, and patterns.
  • Some experience developing software in Linux environments, preferably with Python and/or Golang.
  • A customer-minded approach that enables the success of platform users as well as building trust across the organization.
  • A collaborative disposition that allows you to work optimally on and across teams.
  • Six years of systems experience either in operations or development.

Missing some items on the list? That's ok! We still want to talk to you!

How Our Team Works Together

  • We work out of a shared backlog.
  • We pair-program weekly, as it makes sense.
  • We peer-review everything.
  • We do weekly blame-free retros to reinforce what’s going well and surface what’s not, so we can improve.

Pay

The base salary range for this role is between $167,300 USD and $242,600 USD, with a primary target range of $195,000 to $225,000 based on experience, skills, and market data. This role also includes bonus eligibility and equity.

Benefits

  • Unlimited PTO (modeled and encouraged).
  • Work location flexibility.
  • Up to 24 weeks of parental leave.
  • Excellent health benefits.

Schedule

This role is remote within the United States.

Similar jobs