Data Center Site Reliability Engineer
Synopsys Inc · Canonsburg, PA · 2 wk ago
Information TechnologyFull-time
About the role
Serve as the primary on-site technical resource supporting the Canonsburg data center and coordinating day-to-day operational needs with global engineering teams.
Responsibilities
- Perform Linux systems administration tasks, including basic troubleshooting, log analysis, remote access support, and service management to keep critical systems running reliably.
- Install, rack, cable, relocate, and decommission physical server and networking equipment while maintaining clean, organized rack layouts and labeling standards.
- Maintain accurate infrastructure documentation and digital asset records in DCIM tools (e.g., Sunbird or similar), ensuring inventory, connectivity, and capacity data stays current.
- Coordinate and oversee third-party vendors performing maintenance and infrastructure work, confirming scope, access requirements, safety practices, and completion criteria.
- Monitor data center health indicators such as power, cooling, rack capacity, and environmental conditions, escalating risks and initiating corrective actions as needed.
- Respond to operational incidents as part of a shared on-call rotation, meeting established response SLAs and driving issues through resolution and follow-up.
Impact
- Keep mission-critical engineering systems dependable by reducing downtime and restoring service quickly when issues arise.
- Improve day-to-day operational confidence through accurate asset records and documentation that make capacity, ownership, and change planning clear.
- Accelerate infrastructure deployments by ensuring on-site execution is timely, consistent, and aligned with global engineering standards.
- Reduce operational risk by identifying early warning signs in power, cooling, and environmental conditions and driving corrective actions before they become incidents.
- Strengthen vendor outcomes by ensuring work is properly scoped, safely executed, and fully completed with clear validation and follow-through.
- Increase cross-team effectiveness by serving as a reliable on-site partner who communicates clearly, escalates appropriately, and closes loops after changes and incidents.
- Support successful data center integration and modernization by helping standardize processes and stabilizing operations during periods of change.
Requirements
- You have experience administering Linux systems (Red Hat, Rocky Linux, Ubuntu, or similar) and can navigate common operational tasks with confidence.
- You bring hands-on familiarity working in a physical enterprise data center environment, where safety, precision, and process matter.
- You have a working understanding of server hardware, rack infrastructure, structured cabling, power distribution, and cooling fundamentals.
- You bring strong troubleshooting and problem-solving habits, including the ability to stay calm, prioritize effectively, and drive issues to resolution.
- You have experience coordinating with third-party vendors and service providers and can ensure work is completed to scope and standard.
- You are able to work independently as the primary on-site technical resource and communicate clearly with remote engineering partners.
Skills
- Exposure to DCIM platforms (e.g., Sunbird).
- Experience with network/storage/virtualization environments.
- Scripting and automation with Bash, Python, or PowerShell.
About You
- You are the kind of person who takes ownership end-to-end, following through until the issue is fully resolved and the next steps are clear to everyone involved.
- You approach problems methodically, separating symptoms from root causes and validating changes before and after you act.
- You communicate with precision, tailoring updates to the audience and escalating early when risk, safety, or SLA impact is on the line.
- You stay organized in fast-moving environments, keeping documentation, labels, and records accurate so others can operate confidently after you.
- You collaborate smoothly across teams and vendors, setting expectations upfront and ensuring work is completed safely, cleanly, and to standard.
Team
The Data Center Operations team supports reliable, secure day-to-day execution across Synopsys’ physical infrastructure footprint while partnering closely with global network, storage, security, and infrastructure engineering teams. This role serves as the primary on-site operator for the Canonsburg data center, ensuring operational excellence and enabling modernization and integration efforts.
Benefits
We offer a comprehensive range of health, wellness, and financial benefits to cater to your needs. Our total rewards include both monetary and non-monetary offerings.