Jobs · Washington

Staff Software Engineer - Fleet Management

Nscale · Bellevue, WA · 2 wk ago
Hybrid$220k–$320k/yrFull-time

Nscale is the GPU cloud engineered for AI, providing cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. We enable AI-focused companies to achieve superior results by reducing the complexity of AI development, supporting strategic business outcomes such as cost management, rapid innovation, and environmental responsibility. Our culture thrives on relentless innovation, ownership, and accountability, where every team member drives excellence and urgency.

About the role

Nscale is hiring a Staff Software Engineer to build Fleet Manager—the workflow automation platform that provisions, tests, and remediates GPU nodes and network switches at scale. This role sits at the intersection of distributed systems, infrastructure automation, and physical hardware. You'll own domain-level architecture within Fleet Manager: Python-based systems that manage the entire operational lifecycle of our compute infrastructure, from initial device enrollment through multi-day burn-in testing to ongoing health monitoring and automated remediation.

What you'll work on

  • Device provisioning and enrollment: automation that takes bare-metal GPU nodes and network switches from first power-on to production-ready—BMC configuration, DHCP reservations, and provisioning state machines.
  • Burn-in and validation: multi-day testing workflows that qualify hardware before it enters, and re-enters, the fleet.
  • Workflow orchestration: durable, event-driven state machines that span multiple days, survive crashes, resume from checkpoints, support human-in-the-loop approval gates, and let thousands of concurrent idempotent workflows run without stepping on each other.
  • GPU health monitoring and self-healing: detection, diagnosis, and automated remediation workflows that keep nodes healthy in production.
  • Network configuration: switch lifecycle automation and network state management across the fleet.
  • Integrations: keeping Fleet Manager consistent with datacenter inventory tooling (DCIM, NetBox), bare-metal provisioning systems (MAAS, Ironic, IPMI), credential stores, and monitoring infrastructure.
  • Observability: structured logging, metrics, distributed tracing, and tooling that lets operators troubleshoot effectively.

Responsibilities

  • Domain-level technical direction: Own the architecture for a major Fleet Manager domain—such as provisioning, validation, or remediation—influencing engineers across the team and adjacent squads.
  • Design and build production-grade automation: Implement device provisioning, burn-in testing, network configuration, and hardware health validation workflows in Python.
  • Engineer for reliability and auditability: Treat idempotency, resumability, checkpointing, retries, replay, and failure handling as first-class design concerns.
  • Integrate broadly: Connect Fleet Manager with datacenter infrastructure management systems, cloud orchestration platforms, and bare-metal provisioning tools.
  • Create leverage through standards: Establish shared patterns, libraries, conventions, and operational runbooks that other engineers build on.
  • Own production outcomes: Operate what you build with strong observability, alerting, incident response, and day-2 operational discipline.
  • Mentor and influence: Raise the bar through design reviews, implementation guidance, and operational best practices.
  • Use AI to accelerate delivery while maintaining architectural coherence.

Requirements

  • Extensive experience designing, building, and operating distributed systems in production, ideally in infrastructure automation, workflow tooling, or platform engineering.
  • Strong proficiency in Python—Fleet Manager is built entirely in Python.
  • Strong understanding of event-driven and workflow architecture, including reliable delivery, idempotency, retries, replay, and failure handling.
  • Track record of delivering automation systems from ambiguous requirements to production, with hands-on day-2 operations experience (monitoring, incident response, performance optimization).
  • Proven ability to lead ambiguous technical work across team boundaries and drive domain-level delivery through influence rather than formal authority.
  • Driven by building distributed systems at scale, infrastructure reliability, scalability, security, and continuous improvement.
  • Use AI tools like Claude or Cursor as a core part of your development workflow to create leverage, increase quality, and accelerate delivery.
  • Excellent communication skills to build consensus with stakeholders, both internally and externally, in a fast-paced, high-agency environment.

Preferred

  • Experience with workflow orchestration tools like Temporal, Airflow, Prefect, or similar.
  • Hands-on experience with infrastructure tooling: DCIMs, NetBox, OpenStack, or ERP systems.
  • Bare-metal provisioning and automation: MAAS, Ironic, IPMI, PXE boot, or network automation.
  • Experience building hardware lifecycle automation: provisioning, validation, testing, or remediation workflows.
  • GPU infrastructure experience: health monitoring, burn-in testing, or cluster management.
  • HPC and networking: datacenter topology, high-performance interconnects (InfiniBand, RoCE).
  • Deep knowledge of Kubernetes, Infrastructure as Code (Terraform, Pulumi), AWS, and GCP.
  • Open-source contributions in infrastructure automation or cloud-native tooling.

Pay

Salary Range: $220,000—$320,000 USD

Benefits

  • Medical, dental, and vision coverage.
  • Flexible paid time off.
  • Parental leave.
  • Retirement plan participation.
  • Bonus, equity, and/or commission programs (role-dependent).

Similar jobs