Principal Site Reliability Engineer
About Us
At Synopsys, we drive the innovations that shape the way we live and connect. Our technology is central to the Era of Pervasive Intelligence, from self-driving cars to learning machines. We lead in chip design, verification, and IP integration, empowering the creation of high-performance silicon chips and software content. Join us to transform the future through continuous technological innovation.
About You
You thrive on solving challenges in a large-scale HPC environment and are passionate about creating scalable processes. You enjoy working collaboratively across time zones and cultures, with excellent problem-solving skills and a proactive approach to issues. You are excited about joining an innovative team that values continuous learning, great leadership, and being part of a growing organization.
Responsibilities
- Apply SRE practices to identify, monitor, communicate, and resolve issues in the environment, while collaborating with internal teams and customers on post-mortem analysis to deliver root cause insights.
- Follow up on reported issues and develop procedures to prevent similar occurrences.
- Review current processes and transform them into scalable solutions.
- Debug OS and engineering issues within our provided Linux environment.
- Collaborate on internal projects across different time zones and teams.
- Follow up with customers and hand over tasks/issues with team members to utilize time zones efficiently.
Impact
- Enhance the reliability and performance of our engineering environment.
- Streamline processes to ensure scalability and efficiency.
- Resolve complex OS and engineering issues, contributing to smoother operations.
- Drive successful project outcomes through effective collaboration across time zones.
- Improve customer satisfaction by addressing and resolving issues promptly.
- Foster SRE practices within multifunctional teams and identify gaps for resolution.
Requirements
- 10+ years of SRE processes and related skills.
- Capability to understand complex engineering implementations and their interdependencies for troubleshooting.
- Deep knowledge of Linux distributions (CentOS, RedHat, Ubuntu, SuSE).
- Deep knowledge of virtualization and containerization technologies.
- Extensive knowledge of storage solutions, including network storage and associated protocols.
- Good experience in network technologies.
- Good experience in load sharing facilities such as LSF, Slurm, and various workload scheduling technologies.
- Strong interpersonal, communication, and leadership skills.
Who You Are
- Part of a global team supporting one of the biggest scaled environments, including multiple HPC clusters, high-performance storage, large-scale private cloud implementations, and GPU clusters for HPC/GenAI workloads.
- Challenged by working with the latest state-of-the-art technologies.
- Excited to be part of one of the biggest private clouds in the world.
- Committed to embracing and implementing SRE best practices.
- Able to monitor and comprehend complex environments.
- Skilled at breaking down complex issues into relevant areas and independently coordinating follow-ups with internal teams.
- A good communicator with strong interpersonal skills.
- A proactive problem solver with a keen eye for detail.
- A collaborative team player who thrives in a global, intercultural environment.
- Adept at multitasking and managing multiple priorities effectively.
- Self-motivated and capable of working independently.
- Passionate about continuous learning and professional development.
The Team
You will be part of the Platform Team at Synopsys, a dynamic group dedicated to driving technological innovation. Our team works on cutting-edge projects, ensuring the reliability and performance of our engineering environment. We collaborate across time zones and cultures, leveraging diverse perspectives to achieve our goals.