Jobs · Arizona

Principal Software Engineer - Infrastructure Automation

Early Warning · Scottsdale, AZ · 3 days ago
$173k–$230k/yrFull-time

About the role

The Principal Software Engineer - Infrastructure Automation is responsible for the architecture, strategy, and engineering direction of the infrastructure-as-code platform used to provision and manage infrastructure across Early Warning. This role operates with minimal guidance, identifying enterprise-level infrastructure automation problems and opportunities, defining the target solution and execution path, resolving ambiguous cross-domain decisions, and driving delivery through measurable adoption and outcomes.

Responsibilities

  • Identify and resolve enterprise infrastructure automation problems and opportunities, define the target solution and execution path, resolve cross-domain tradeoffs, and drive delivery through measurable adoption and outcomes.
  • Establish software engineering standards for infrastructure modules, policy, and tooling, including versioned interfaces, testing, code review, release management, deprecation, and automated regression coverage.
  • Own Terraform as the enterprise default for infrastructure provisioning across cloud and on-premises estates, including module taxonomy, ownership, private registry strategy, and adoption standards.
  • Define and deliver opinionated Terraform modules for AWS services and common service patterns that provide secure, compliant, observable defaults for encryption, logging, tagging, networking, and IAM.
  • Establish the enterprise override and exception model so legitimate deviations remain visible, reviewable, policy-checked, time-bound, and auditable.
  • Architect the Terraform module test and release system, including native Terraform tests, Terratest, tflint, static analysis, semantic versioning, signed artifacts, and policy gates on plan output.
  • Own the HCP Terraform or Terraform Enterprise operating model, including workspaces, run tasks, policy sets, protected state, drift detection, and plan/apply governance independent of the CI system that initiates a run.
  • Own the enterprise Ansible platform architecture, including collections, roles, execution environments, Molecule and lint standards, and Ansible Automation Platform/AWX operations.
  • Define the configuration-as-code architecture for network infrastructure, including structured configuration as source of truth, safe rollback, virtual topology and reachability validation, post-change verification, compliance, and drift detection.
  • Own the enterprise policy-as-code architecture and Rego policy library for Terraform plans, Ansible constraints, and Kubernetes admission; ensure policy is tested, versioned, released, and distributed as a governed product.
  • Design programmatic controls for segregation of duties, protected state, change review, and break-glass access with automatic evidence capture and time-boxed expiration.
  • Eliminate static credentials from infrastructure change workflows through OIDC-federated, short-lived IAM roles and dynamic secrets for cloud, database, and network credentials.
  • Define the infrastructure automation control framework and map enforcement points to PCI DSS, SOX ITGC, NYDFS Part 500, FFIEC, and other applicable requirements; convert manual controls into continuously evidenced automated controls.
  • Serve as the technical authority for infrastructure-as-code controls in internal audits, external audits, and regulatory examinations, and author enterprise engineering standards and control narratives.
  • Architect self-service capabilities for module discovery, scaffolding, exception workflows, compliance dashboards, and APIs that expose module, policy, and compliance state to internal platforms.
  • Prototype and de-risk the most complex infrastructure automation and guardrail problems; mentor senior engineers and raise the architecture, design, and code-review bar across engineering.
  • Partner with CI/CD, Security, Network Engineering, AIOps, and observability leaders to integrate infrastructure modules, controls, telemetry, and governance across shared engineering platforms.
  • Define and report enterprise measures of platform effectiveness, including module adoption, policy pass rate, plan/apply success, control coverage, and time to remediate drift and compliance violations.
  • Support the company's commitment to risk management and protecting the integrity and confidentiality of systems and data.

Requirements

Education and/or experience typically obtained through a bachelor's degree in computer science, engineering, or a related technical field. Ten or more years of related experience in software, platform, cloud, or infrastructure engineering, including significant experience leading enterprise-scale technical architecture. Demonstrated excellence in software engineering and distributed-systems design, including rigorous reasoning about interfaces, state, failure modes, idempotency, and blast radius. Demonstrated ability to operate with very little guidance, proactively identify unstructured technical problems or opportunities, define the solution and execution strategy, and drive complex work from concept through implementation and organizational adoption. Demonstrated customer-first leadership at enterprise scale, balancing customer and developer outcomes with risk, reliability, and control requirements. Demonstrated ability to model disagree and commit: challenge assumptions with evidence, drive timely decisions, and align the organization behind the final direction. Strong production programming experience in Python or Go building tools, controllers, services, APIs, or custom scanners - not solely HCL and YAML. Demonstrated experience architecting and operating platform capabilities used by multiple engineering organizations, including testing, release management, reliability, and operational ownership. Deep Terraform expertise at enterprise scale, including module composition, testing, private registries, state governance, version strategy, and plan/apply controls. Deep experience with Ansible beyond playbook development, including collections, execution environments, Molecule testing, linting, and Ansible Automation Platform/AWX operations. Fluency with OPA/Rego or a comparable policy language, including authoring, testing, versioning, and operating policy-as-code libraries. Deep knowledge of AWS infrastructure and security, particularly IAM, federated identity, short-lived credentials, networking, encryption, and secure configuration patterns. Working knowledge of Kubernetes sufficient to design and evaluate infrastructure, identity, and admission-control policies. Demonstrated experience embedding security, regulatory, and change-management controls into automated infrastructure workflows in financial services or a comparably regulated industry. Demonstrated ability to influence enterprise architecture and engineering standards across organizational boundaries without direct authority. Exceptional communication skills, including architecture decisions, engineering standards, executive communication, audit narratives, and regulatory control documentation.

Qualifications

Education and/or experience typically obtained through a bachelor's degree in computer science, engineering, or a related technical field. Ten or more years of related experience in software, platform, cloud, or infrastructure engineering, including significant experience leading enterprise-scale technical architecture. Demonstrated excellence in software engineering and distributed-systems design, including rigorous reasoning about interfaces, state, failure modes, idempotency, and blast radius. Demonstrated ability to operate with very little guidance, proactively identify unstructured technical problems or opportunities, define the solution and execution strategy, and drive complex work from concept through implementation and organizational adoption. Demonstrated customer-first leadership at enterprise scale, balancing customer and developer outcomes with risk, reliability, and control requirements. Demonstrated ability to model disagree and commit: challenge assumptions with evidence, drive timely decisions, and align the organization behind the final direction. Strong production programming experience in Python or Go building tools, controllers, services, APIs, or custom scanners - not solely HCL and YAML. Demonstrated experience architecting and operating platform capabilities used by multiple engineering organizations, including testing, release management, reliability, and operational ownership. Deep Terraform expertise at enterprise scale, including module composition, testing, private registries, state governance, version strategy, and plan/apply controls. Deep experience with Ansible beyond playbook development, including collections, execution environments, Molecule testing, linting, and Ansible Automation Platform/AWX operations. Fluency with OPA/Rego or a comparable policy language, including authoring, testing, versioning, and operating policy-as-code libraries. Deep knowledge of AWS infrastructure and security, particularly IAM, federated identity, short-lived credentials, networking, encryption, and secure configuration patterns. Working knowledge of Kubernetes sufficient to design and evaluate infrastructure, identity, and admission-control policies. Demonstrated experience embedding security, regulatory, and change-management controls into automated infrastructure workflows in financial services or a comparably regulated industry. Demonstrated ability to influence enterprise architecture and engineering standards across organizational boundaries without direct authority. Exceptional communication skills, including architecture decisions, engineering standards, executive communication, audit narratives, and regulatory control documentation.

Skills

Strong production programming experience in Python or Go building tools, controllers, services, APIs, or custom scanners - not solely HCL and YAML. Demonstrated experience architecting and operating platform capabilities used by multiple engineering organizations, including testing, release management, reliability, and operational ownership. Deep Terraform expertise at enterprise scale, including module composition, testing, private registries, state governance, version strategy, and plan/apply controls. Deep experience with Ansible beyond playbook development, including collections, execution environments, Molecule testing, linting, and Ansible Automation Platform/AWX operations. Fluency with OPA/Rego or a comparable policy language, including authoring, testing, versioning, and operating policy-as-code libraries. Deep knowledge of AWS infrastructure and security, particularly IAM, federated identity, short-lived credentials, networking, encryption, and secure configuration patterns. Working knowledge of Kubernetes sufficient to design and evaluate infrastructure, identity, and admission-control policies. Demonstrated experience embedding security, regulatory, and change-management controls into automated infrastructure workflows in financial services or a comparably regulated industry. Exceptional communication skills, including architecture decisions, engineering standards, executive communication, audit narratives, and regulatory control documentation.

Benefits

Healthcare Coverage - Competitive medical (PPO/HDHP), dental, and vision plans as well as company contributions to your Health Savings Account (HSA) or pre-tax savings through flexible spending accounts (FSA) for commuting, health & dependent care expenses. 401(k) Retirement Plan - Featuring a 100% Company Safe Harbor Match on your first 6% deferral immediately upon eligibility. Paid Time Off - Flexible Time Off for Exempt (salaried) employees, as well as generous PTO for Non-Exempt (hourly) employees, plus 11 paid company holidays and a paid volunteer day. 12 weeks of Paid Parental Leave Maven Family Planning - provides support through your Parenting journey including egg freezing, fertility, adoption, surrogacy, pregnancy, postpartum, early pediatrics, and returning to work. And SO much more! We continue to enhance our program, so be sure to check our Benefits page.

Pay

The Base Pay Scale For This Position In Phoenix, AZ/ Chicago, IL in USD per year is: $173,000 - $230,000. Additionally, candidates are eligible for a discretionary incentive plan and benefits. This pay scale is subject to change and is not necessarily reflective of actual compensation that may be earned, nor a promise of any specific pay for any specific candidate, which is always dependent on legitimate factors considered at the time of job offer. Early Warning Services takes into consideration a variety of factors when determining a competitive salary offer, including, but not limited to, the job scope, market rates and geographic location of a position, candidate’s education, experience, training, and specialized skills or certification(s) in relation to the job requirements and compared with internal equity (peers).

Similar jobs