Grid Operations Architect
Siemens Digital Industries Software · Wilsonville, OR · Yesterday
Hybrid$130k–$233k/yrFull-time
About the role
We are a leading global software company dedicated to the world of computer aided design, 3D modeling and simulation— helping innovative global manufacturers design better products, faster! With the resources of a large company, and the energy of a software start-up, we have fun together while creating a world class software portfolio. Our culture encourages creativity, welcomes fresh thinking, and focuses on growth, so our people, our business, and our customers can achieve their full potential.
Responsibilities
- Own the multiyear architectural roadmap for scaling grid compute capacity from 263K+ cores to 1M+ cores across 18+ globally distributed grids, consolidating across sites while maintaining availability and performance.
- Design the phased migration plan from AGE to Accelerator's hierarchical scheduling platform.
- Develop consolidation strategies to reduce operational complexity rationalizing 4 scheduler platforms toward a unified Accelerator based model, standardizing tooling, and harmonizing operational policies across regions.
- Lead capacity planning and forecasting models that project compute demand growth using linear, exponential, and flat percentage projections. Translate business requirements into infrastructure investment plans aligned with the 80% utilization ceiling and 72% procurement trigger governance model.
- Architect hybrid and cloud burst strategies (e.g., NavOps ) to augment on premises capacity for peak demand and elastic workloads working with A&S and Networking teams while respecting the latency and data locality requirements of EDA workloads.
- Design fault - tolerant, self healing grid architectures that minimize downtime and reduce manual intervention critical for enabling 24x5 follow the sun coverage without proportional headcount growth.
- Establish architectural standards, reference designs, and best practices for grid deployment, configuration, and lifecycle management across all 18+ grids and 9+ clusters.
- Design and drive the evolution of the grid automation and tooling ecosystem — modernizing legacy Perl/Shell-based operational tools toward scalable, maintainable platforms (Python, APIs, infrastructure-as-code, GitLab CI/CD).
- Target: >50% reduction in manual operational toil across 8 identified operational areas.
- Architect the observability and analytics stack (Elastic Stack with geo-redundant instances across Americas and Asia, Grafana, Power BI) to provide real-time visibility into grid performance, availability, utilization, and capacity across all grids.
- Develop AI-assisted alerting and auto-classification capabilities leveraging Elastic AI Agent and ServiceNow API integration to move from reactive monitoring to predictive operations.
- Evolve SLA frameworks and operational KPIs from current state to improve reporting, performance: 99.97% availability per quarter, >97% SLO compliance, 30-minute MTTR, job wait time tracking at P50 and P95.
- Partner with engineering leadership, R&D teams, and business stakeholders to understand compute demand trends, workload characteristics, and evolving requirements.
- Collaborate with A&S team, Infrastructure Technology Verticals — networking, storage, virtualization, and data center operations — to ensure end-to-end infrastructure coherence and performance.
- Drive continuous improvement through data-driven analysis of incidents, performance trends, and utilization patterns using Elastic dashboards and capacity modelling tools (Griddash).
- Design self-service onboarding and provisioning workflows (ServiceNow catalogue items) to reduce manual provisioning overhead and improve time-to-compute for engineering teams.
- Ensure software license compliance across all grid engines and associated tooling — particularly critical during the scheduler consolidation from 3 license agreements to 1.
- Contribute to disaster recovery and business continuity planning for grid services.
- Participate in daily operations standups, weekly staff meetings, and bi-weekly cadences with key customer teams. Provide architectural input on operational issues and escalations.
Qualifications
- Bachelor's degree (or equivalent) in Computer Science, Information Technology, Engineering, or related field. Master's degree preferred.
- 15+ years of progressive experience in systems engineering, infrastructure architecture, or HPC/grid computing environments, with at least 5 years in an architecture or technical leadership role.
- Deep expertise with distributed computing platforms — hands-on experience architecting and operating at least two of: Altair Grid Engine, Altair Accelerator, IBM LSF, SLURM, or equivalent workload schedulers.
- Demonstrated experience scaling compute infrastructure in large, multi-site, enterprise-wide environments (100K+ cores).
- Experience with scheduler migrations or platform consolidation at production scale.
- Strong systems-level knowledge of Linux (RHEL/CentOS), including kernel tuning, performance optimization, and large-scale fleet management.
- Proficiency in infrastructure automation and tooling — Python, Perl, Shell scripting — with experience in modern approaches (Ansible, Terraform, infrastructure-as-code, CI/CD pipelines).
- Solid understanding of networking (TCP/IP, DNS, NFS/GPFS, InfiniBand), storage architectures, and virtualisation as they relate to HPC and grid computing.
- Experience in cloud computing platforms (AWS, Azure, GCP, OCI) and hybrid cloud/cloud-burst architectures for HPC workloads (e.g., NavOps, Cycle Cloud).
- Proven ability to develop and communicate multi-year technical roadmaps and translate business requirements into architectural decisions with clear investment justification.
- Strong project management, prioritization, and stakeholder management skills with experience presenting to senior leadership and executive audiences.
- Excellent verbal and written communication skills with the ability to present complex technical concepts to both technical and non-technical audiences.