Jobs · Oregon

Grid Operations Architect

Siemens Digital Industries Software · Wilsonville, OR · Yesterday
Hybrid$130k–$233k/yrFull-time

About the role

We are a leading global software company dedicated to the world of computer aided design, 3D modeling and simulation— helping innovative global manufacturers design better products, faster! With the resources of a large company, and the energy of a software start-up, we have fun together while creating a world class software portfolio. Our culture encourages creativity, welcomes fresh thinking, and focuses on growth, so our people, our business, and our customers can achieve their full potential.

Responsibilities

  • Own the multiyear architectural roadmap for scaling grid compute capacity from 263K+ cores to 1M+ cores across 18+ globally distributed grids, consolidating across sites while maintaining availability and performance.
  • Design the phased migration plan from AGE to Accelerator's hierarchical scheduling platform.
  • Develop consolidation strategies to reduce operational complexity rationalizing 4 scheduler platforms toward a unified Accelerator based model, standardizing tooling, and harmonizing operational policies across regions.
  • Lead capacity planning and forecasting models that project compute demand growth using linear, exponential, and flat percentage projections. Translate business requirements into infrastructure investment plans aligned with the 80% utilization ceiling and 72% procurement trigger governance model.
  • Architect hybrid and cloud burst strategies (e.g., NavOps ) to augment on premises capacity for peak demand and elastic workloads working with A&S and Networking teams while respecting the latency and data locality requirements of EDA workloads.
  • Design fault - tolerant, self healing grid architectures that minimize downtime and reduce manual intervention critical for enabling 24x5 follow the sun coverage without proportional headcount growth.
  • Establish architectural standards, reference designs, and best practices for grid deployment, configuration, and lifecycle management across all 18+ grids and 9+ clusters.
  • Design and drive the evolution of the grid automation and tooling ecosystem — modernizing legacy Perl/Shell-based operational tools toward scalable, maintainable platforms (Python, APIs, infrastructure-as-code, GitLab CI/CD).
  • Target: >50% reduction in manual operational toil across 8 identified operational areas.
  • Architect the observability and analytics stack (Elastic Stack with geo-redundant instances across Americas and Asia, Grafana, Power BI) to provide real-time visibility into grid performance, availability, utilization, and capacity across all grids.
  • Develop AI-assisted alerting and auto-classification capabilities leveraging Elastic AI Agent and ServiceNow API integration to move from reactive monitoring to predictive operations.
  • Evolve SLA frameworks and operational KPIs from current state to improve reporting, performance: 99.97% availability per quarter, >97% SLO compliance, 30-minute MTTR, job wait time tracking at P50 and P95.
  • Partner with engineering leadership, R&D teams, and business stakeholders to understand compute demand trends, workload characteristics, and evolving requirements.
  • Collaborate with A&S team, Infrastructure Technology Verticals — networking, storage, virtualization, and data center operations — to ensure end-to-end infrastructure coherence and performance.
  • Drive continuous improvement through data-driven analysis of incidents, performance trends, and utilization patterns using Elastic dashboards and capacity modelling tools (Griddash).
  • Design self-service onboarding and provisioning workflows (ServiceNow catalogue items) to reduce manual provisioning overhead and improve time-to-compute for engineering teams.
  • Ensure software license compliance across all grid engines and associated tooling — particularly critical during the scheduler consolidation from 3 license agreements to 1.
  • Contribute to disaster recovery and business continuity planning for grid services.
  • Participate in daily operations standups, weekly staff meetings, and bi-weekly cadences with key customer teams. Provide architectural input on operational issues and escalations.

Qualifications

  • Bachelor's degree (or equivalent) in Computer Science, Information Technology, Engineering, or related field. Master's degree preferred.
  • 15+ years of progressive experience in systems engineering, infrastructure architecture, or HPC/grid computing environments, with at least 5 years in an architecture or technical leadership role.
  • Deep expertise with distributed computing platforms — hands-on experience architecting and operating at least two of: Altair Grid Engine, Altair Accelerator, IBM LSF, SLURM, or equivalent workload schedulers.
  • Demonstrated experience scaling compute infrastructure in large, multi-site, enterprise-wide environments (100K+ cores).
  • Experience with scheduler migrations or platform consolidation at production scale.
  • Strong systems-level knowledge of Linux (RHEL/CentOS), including kernel tuning, performance optimization, and large-scale fleet management.
  • Proficiency in infrastructure automation and tooling — Python, Perl, Shell scripting — with experience in modern approaches (Ansible, Terraform, infrastructure-as-code, CI/CD pipelines).
  • Solid understanding of networking (TCP/IP, DNS, NFS/GPFS, InfiniBand), storage architectures, and virtualisation as they relate to HPC and grid computing.
  • Experience in cloud computing platforms (AWS, Azure, GCP, OCI) and hybrid cloud/cloud-burst architectures for HPC workloads (e.g., NavOps, Cycle Cloud).
  • Proven ability to develop and communicate multi-year technical roadmaps and translate business requirements into architectural decisions with clear investment justification.
  • Strong project management, prioritization, and stakeholder management skills with experience presenting to senior leadership and executive audiences.
  • Excellent verbal and written communication skills with the ability to present complex technical concepts to both technical and non-technical audiences.

Similar jobs