Sr. Staff Site Reliability Engineer
About the role
At Synopsys, we drive and shape the way we live and connect. Our technology is central to the Era of Pervasive Intelligence, from self-driving cars to learning machines. We lead in chip design, verification, and IP integration, empowering the creation of high-performance silicon chips and software content. Join us to transform the future through continuous technological innovation.
You will be part of a global team supporting one of the biggest scaled environments that includes multiple HPC clusters, high-performance storage, large-scale private cloud implementation, as well as GPU clusters for HPC/GenAI workloads. You’ll challenge yourself with the latest state-of-the-art technologies and be part of one of the biggest private clouds in the world.
Responsibilities
- Focus on automation and process improvements.
- Enhance our virtualization and containerization infrastructure.
- Troubleshoot operating system and engineering issues within our Linux environment.
- Collaborate on internal projects across various time zones and teams.
- Coordinate with customers and transfer tasks/issues to team members to optimize time zone utilization.
The impact you will have:
- Improve the reliability and performance of our engineering environment.
- Scale and manage our engineering environment more effectively.
- Streamline processes to enhance scalability and efficiency.
- Address complex operating system and engineering challenges to facilitate smoother operations.
- Ensure successful project outcomes through effective collaboration across different time zones.
- Improve customer satisfaction by addressing and resolving issues promptly.
Requirements
- Proficient in Linux operating systems, including RHEL, SLES, and Ubuntu.
- Ability to comprehend complex engineering implementations and their interdependencies for effective troubleshooting.
- Experienced with automation platforms such as Ansible, Tower, or AWX.
- Familiar with various architectures, including x86_64, aarch64, and ppc64.
- Extensive knowledge of virtualization and containerization technologies.
- Solid understanding of networking fundamentals.
- Well-versed in storage solutions, including local storage and network appliances.
- Strong interpersonal and communication skills.
Qualifications
- Embrace and implement SRE best practices.
- Able to monitor and comprehend complex environments.
- Skilled at breaking down complex issues into relevant areas and independently coordinating follow-ups with internal teams.
- A proactive problem solver with a keen eye for detail.
- A collaborative team player who thrives in a global, intercultural environment.
- Adept at multitasking and managing multiple priorities effectively.
- Self-motivated and capable of working independently.
- Passionate about continuous learning and professional development.
The Team You’ll Be a Part Of
You will be part of the Platform Team at Synopsys, a dynamic group dedicated to driving technological innovation. Our team works on cutting-edge projects, ensuring the reliability and performance of our engineering environment. We collaborate across time zones and cultures, leveraging diverse perspectives to achieve our goals.