Batch Processing System Architect
Job Description
HPE is seeking a Linux Systems and Automation Architect to support and enhance the Linux-based infrastructure used for ASIC chip development. This infrastructure includes RHEL8 servers (VMs and bare metal), NFS file servers, MariaDB databases, and LSF job scheduling systems, all extensively automated using Ansible and custom scripts.
Key Responsibilities
- Manage and configure LSF Scheduler/Resource Manager and RTM monitoring to optimize ASIC development workloads and collaborate with the ASIC tools team to maintain and evolve LSF configurations.
- Install, update, and support EDA tools from vendors such as Cadence, Synopsys, and Mentor, including managing Flexera Licensing and license file configurations.
- Develop and maintain automation scripts (Ansible, Bash, Python, Perl) for system deployments, updates, configuration management, and LSF environment operations across large-scale Linux environments.
- Administer RHEL Linux systems at scale, including NFS server management, user account and LDAP administration, certificate management, patching, and system hardening.
- Monitor and optimize system performance across network, storage, and compute resources, and assist ASIC teams in investigating and resolving tool-related issues.
- Cook up and maintain comprehensive documentation including operational runbooks, automated process guides, and change management records, while coordinating system events such as patching and planned downtime with business units.
Requirements
- Bachelor's degree in Computer Science, Information Technology, or a related field, or equivalent combination of education and experience.
- Typically, 6+ years of experience managing RHEL Linux systems at scale; RHEL or equivalent Linux certifications strongly preferred.
- Hands-on experience with LSF Scheduler/Resource Manager, RTM monitoring, or similar workload management platforms in an ASIC or EDA environment.
- Proficiency in scripting languages including Python, Bash, and Perl for automation and systems management.
- Experience with automation and configuration management tools such as Ansible.
- Familiarity with ASIC EDA tools (Cadence, Synopsys, Mentor) and Flexera Licensing.
- Experience managing NFS servers, Linux patching (yum, RPMs), LDAP authentication, and system hardening tools (SaltStack, VMware Aria, or similar).
- Experience with incident and change management platforms such as ServiceNow, including provisioning and decommissioning workflows.
Skills
- Strong communication and collaboration skills with the ability to work effectively across cross-functional teams including ASIC engineers, infrastructure teams, and management.
- Ability to manage multiple priorities and deliver results in a fast-paced, high-performance computing environment.
- Strong documentation skills for creating runbooks, operational workflows, and process guides.
- Knowledge of NFS storage infrastructure, network troubleshooting, and Linux performance optimization.
- Familiarity with remote access tools such as NoMachine (NX) and graphical desktop environments (GNOME/KDE/MATE).
- Understanding of cloud infrastructure coordination, including architecture design collaboration and system provisioning processes.
What We Can Offer You
- Health & Wellbeing
- Personal & Professional Development
- Unconditional Inclusion
Job Information
Job Level: TCP_04
The expected salary/wage range for this position is provided below. Actual offer may vary from this range based upon geographic location, work experience, education/training, and/or skill level.
- United States of America: Annual Salary USD 105,500 - 213,500 in California // 92,700 - 213,500 in Texas
Information about employee benefits offered in the US can be found at https://myhperewards.com/main/new-hire-enrollment.html.
HPE is an Equal Employment Opportunity/ Veterans/ Disabled/LGBT employer. We do not discriminate on the basis of race, gender, or any other protected category, and all decisions we make are made on the basis of qualifications, merit, and business need.