Site Reliability Engineer
What You’ll Do
- Designs and deploys small to mid-size or moderately complex solutions to optimize reliability, availability, latency, and performance.
- Integrates knowledge of design, automation, and deployment with expertise in coding to improve service reliability for existing or new systems and adapts for regions, countries, or customers.
- Designs and tests high availability and disaster recovery measures for our services to ensure automation is improving reliability, scalability, and velocity.
- Forecasts and builds reports to determine at what point resources will be at capacity.
- Designs and implements tools that provide visibility into performance and reliability of our infrastructure.
- Builds automated platforms.
- Monitors the environment and works with Developers and Ops to identify problems and develop monitoring tools that provide visibility into performance and reliability, serves as on-call SRE, leads post mortems, and writes root cause analysis.
- Builds and ensures security controls are in place in regards to architectural design, collaborates with security in designing or providing input to security controls, and may actively contribute in security incident response.
Responsibilities
- Evaluates the scalability, resiliency, performance, and security properties and techniques used in production environments.
- Supports the uptime of production services through an On-Call rotation, which includes monitoring and alerting to meet internal Service Level Objectives (SLOs) and customer-facing Service Level Agreements (SLAs).
- Ensures reliable incident processes by conducting Disaster Recovery drills.
- Improves reliability through incident management by investigating incidents, implementing remediation strategies, and learning from past incidents to make improvements.
- Determines the reliability and security requirements of components and systems to meet the reliability objectives of the company, customers, and any relevant governmental agencies.
- Reduces operational expenses through automation, by identifying and mitigating failure points, and automating repetitive and resource-intensive tasks.
- Develops new acceleration techniques and analytical tools to ensure the early identification of potential issues with new products, packaging, processes, and overall product reliability.
- Ensures a reliable and scalable network, manages network / cloud infrastructure and storage systems supporting business operations, and responds to planned maintenance, real-time outages, and issues.
- Plans, designs and implements local and wide-area network solutions between multiple platforms and protocols.
Minimum Qualifications
- Bachelors + 7 years of related experience, or Masters + 4 years of related experience, or PhD + 1 year of related experience.
- Requires solid conceptual and practical knowledge in primary technical job family and knowledge of related technical job families; has worked with a range of technologies.
- 7+ yrs experienced professional using best practices and knowledge of internal or external business issues to improve products or services.
- Works independently, but receives minimal guidance and direction from leader then determines best approach to accomplish work.
- Acts as a resource for colleagues with less experience.
- Understands project and/or department needs and establishes relationships with appropriate cross-functional stakeholders to gather input, collect information, and complete work steps.
Pay & Benefits
The starting salary range for this position is $165,000.00 to $241,400.00 and reflects the projected salary range for new hires in U.S. and/or Canada locations, not including incentive compensation*, equity, or benefits. Individual pay is determined by the candidate's hiring location, market conditions, job-related skillset, experience, qualifications, education, certifications, and/or training.
The applicable full salary ranges for this position, by specific state, are listed below:
- New York City Metro Area: $165,000.00 – $277,600.00
- Non-Metro New York State & Washington State: $146,700.00 – $247,000.00
U.S. employees are offered benefits, subject to Cisco’s plan eligibility rules, which include medical, dental and vision insurance, a 401(k) plan with a Cisco matching contribution, paid parental leave, short and long-term disability coverage, and basic life insurance.
Employees may be eligible to receive grants of Cisco restricted stock units, which vest following continued employment with Cisco for defined periods of time.
U.S. employees are eligible for paid time away as described below, subject to Cisco’s policies:
- 10 paid holidays per full calendar year, plus 1 floating holiday for non-exempt employees
- 1 paid day off for employee’s birthday
- Paid year-end holiday shutdown
- 4 paid days off for personal wellness determined by Cisco
- Non-exempt employees receive 16 days of paid vacation time per full calendar year, accrued at rate of 4.92 hours per pay period for full-time employees
- Exempt employees participate in Cisco’s flexible vacation time off program, which has no defined limit on how much vacation time eligible employees may use (subject to availability and some business limitations)
- 80 hours of sick time off provided on hire date and each January 1st thereafter, and up to 80 hours of unused sick time carried forward from one calendar year to the next
- Additional paid time away may be requested to deal with critical or emergency issues for family members
- Optional 10 paid days per full calendar year to volunteer
For non-sales roles, employees are also eligible to earn annual bonuses subject to Cisco’s policies. Employees on sales plans earn performance-based incentive pay on top of their base salary, which is split between quota and non-quota components, subject to the applicable Cisco plan.