Jobs · Engineering

Senior Site Reliability Engineer, Data Persistence

Wikimedia Foundation · United States · 1 wk ago
RemoteRemoteEngineering$15/hrFull-time

About the role

The Wikimedia Foundation is seeking a Senior Site Reliability Engineer to support and develop the platform serving Wikipedia, one of the world's most popular encyclopedias. The SRE team is responsible for ensuring the reliability and scalability of Wikimedia's global infrastructure.

Responsibilities

  • Perform day-to-day operational and DevOps tasks on Wikimedia’s public facing infrastructure (deployment, maintenance, configuration, troubleshooting)
  • Implement and utilize configuration management and deployment tools (Puppet, Kubernetes)
  • Lead continuous improvement by automating the installation, configuration, and maintenance of services on our platform
  • Work closely with product teams to assist in the architectural design of new services and make them operate at scale
  • Collaborate with a global, cross-functional team in an asynchronous communication environment
  • Mentor peers in areas of technical and operational strength

Requirements

  • 6+ years experience in an SRE/Operations/DevOps role as part of a team
  • Experience with shell and scripting languages (Python, Go, Bash, Ruby; primarily Python)
  • Experience with configuration management tools (Puppet, Ansible)
  • Experience with distributed caching systems and their optimization
  • Experience with package management on Linux systems (Debian)
  • Strong Linux system-level troubleshooting skills
  • Experience automating tasks and processes, identifying process gaps, and finding automation opportunities
  • Strong English language skills (verbal and written)
  • Ability to work independently and effectively in a globally distributed team
  • Experience leading and participating in incident response and post-incident review rituals

Skills

  • Experience with Linux kernel tuning
  • Experience with monitoring, metrics, and logging infrastructure (Prometheus, Grafana)
  • Experience developing or contributing to Free and Open Source software
  • Experience with LAMP stack technologies (PHP/HHVM, memcached/Redis)
  • Experience with defining cross-team Service Level Objectives (SLOs) and their implementation
  • Experience managing backups at scale, preferably using Bacula

Qualifications

  • Experience with other advanced distributed storage and database systems (Cassandra, MariaDB, etc.)
  • Experience operating an on-premise filesystem or object store at scale (OpenStack Swift or Ceph)

Benefits

We offer a competitive salary range for this position, with the anticipated annual pay range for applicants based within the United States being US$ 116,633 to US$ 181,243. For applicants located outside the US, the pay range will be adjusted to the country of hire. We do not ask for or consider salary history and offer a competitive salary based on individual skills, experience, and location.

Pay

The anticipated annual pay range for this position is US$ 116,633 to US$ 181,243 for applicants based within the United States. For applicants located outside the US, the pay range will be adjusted to the country of hire.

Schedule

This is a remote-first role with staff members based in 40+ countries. The SRE team works in a 24/7 on-call rotation, and there may be occasional in-person events and team meetings.

Equal Opportunity Employer

The Wikimedia Foundation is an equal opportunity employer and values diversity. We encourage applications from people with diverse backgrounds and experiences.

Similar jobs