Senior Site Reliability Engineer
About the role
As a Senior Site Reliability Engineer at Proofpoint you will develop a deep understanding of the various services and applications that come together to deliver Proofpoint’s next generation security products. You will contribute to the architecture to improve scalability, operability, service reliability, capacity, and performance. You will be responsible for provisioning, maintaining, and scaling our production services within server farms across multiple, world-wide data centers as well as AWS. We are looking for passion, curiosity, attention to details, taking pride in one’s work, taking ownership, and having ideas and opinions. If you’re an enthusiastic team player who cares about the infrastructure, remains calm in crisis, collaborates cross functionally, loves automation, then we want to talk to you.
Responsibilities
- Build long lasting, effective partnerships across the organization to foster collaboration between Product, Engineering and Operations teams.
- Participate in a 24x7, multi-site production infrastructure powering the Proofpoint services, including deployment, maintenance, troubleshooting, performance tuning, and security.
- Root-cause complex problems and involve multiple stakeholders, network, hardware and software that relate to scaling and performance.
- Ensure proper monitoring, alerting, capacity planning and reporting in the production environment.
- Contribute to the evolving design and architecture of reliable and scalable infrastructure.
- First line of defense during working hours should any alerts or incidents arise.
- Collaborate with product engineering teams to ensure Operations standards are observed, determine resource impacts for upcoming product deployments, and ensure successful product rollouts.
- Participate in an on-call rotation and be willing to jump on escalated issues as needed.
Requirements
- Demonstrable skills and 7+ years’ experience in troubleshooting, and tuning in systems, networking, and cloud services.
- A friendly and collaborative demeanor, working well with cross functional teams.
- Experience with Infrastructure-as-Code using Terraform, Ansible, Cloudformation, etc.
- Experience automating management and operational tasks using Node.Js, Python, or scripting languages.
- Experience with industry-standard foundation technologies such as TCP/IP, HTTP, DNS, and LDAP.
- Experience in management of a large distributed computing environment.
- Experience in management of cloud infrastructure.
- Experience with virtualization – KVM, VMware vSphere, and OpenStack
- Excellent verbal and written communication skills.
- Experience with monitoring and alerting systems.
- Experience with industry-standard operational practices such as change management, incident management, and working in colocation datacenters.
- Experience with public cloud providers such as Amazon EC2 or Microsoft Azure.
- BS or MS in Computer Science, Engineering or related technical discipline, or equivalent experience required.
- US citizenship required
Benefits
- Competitive compensation
- Comprehensive benefits
- Career success on your terms
- Flexible work environment
- Annual wellness and community outreach days
- Always on recognition for your contributions
- Global collaboration and networking opportunities
- Flexible time off
- Comprehensive well-being program with two paid Wellbeing Days and two paid Volunteer Days per year
- Three-week Work from Anywhere option
Pay
- SF Bay Area, New York City Metro Area Base Pay Range 136,200.00 - 214,005.00 USD
- California (excludes SF Bay Area), Colorado, Connecticut, Illinois, Washington DC Metro, Maryland, Massachusetts, New Jersey, Texas, Washington, Virginia, and Alaska Base Pay Range 112,700.00 - 177,100.00 USD
- All other cities and states excluding those listed above Base Pay Range 101,600.00 - 159,720.00 USD
- This role may be eligible for variable compensation and/or equity.