Jobs · Washington

Software Engineering Manager - Fleet Management

Nscale · Bellevue, WA · 2 wk ago
Hybrid$300k–$350k/yrFull-time

Nscale is the GPU cloud engineered for AI, providing cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. We enable AI-focused companies to achieve superior results by reducing the complexity of AI development, bolstering technical capabilities, and supporting strategic business outcomes like cost management, rapid innovation, and environmental responsibility. Our culture thrives on relentless innovation, ownership, and accountability, where every team member drives excellence and urgency.

About the Role

We're hiring a Software Engineering Manager to lead the team building Fleet Manager—the workflow automation platform that provisions, tests, and remediates GPU nodes and network switches at scale. Reporting to the Director of Software Engineering, Fleet Management (based in EMEA), you'll lead engineers building Python-based systems that manage the entire operational lifecycle of our compute infrastructure: device enrollment, burn-in testing, network configuration, GPU health monitoring, and automated remediation.

You'll own delivery and team health end-to-end—planning, execution, hiring, and career development—while partnering closely with Principal and Staff engineers on technical architecture. In our hands-on engineering culture, you'll stay close to the systems through design reviews, incident deep-dives, and code, with people leadership as your primary craft.

Responsibilities

  • Lead and grow the team
    • Hire, onboard, and develop software engineers who thrive in a high-autonomy, high-accountability environment.
    • Coach performance and career growth, set clear expectations, and handle performance management directly and fairly.
    • Own on-call health, handover quality, and a sustainable operational load for the team.
  • Own delivery and execution
    • Turn the Fleet Manager roadmap into executable plans: drive prioritization, manage dependencies and risks, and ship on commitments.
    • Run the team's planning, review, and escalation mechanisms, and remove blockers before they become delays.
    • Communicate status, risks, and trade-offs clearly to stakeholders—including when the news is bad.
  • Drive engineering and operational excellence
    • Uphold engineering standards: code review, testing, CI/CD, incident response, and postmortems that produce durable fixes rather than closed tickets.
    • Own the operational health of the team's services: SLOs, observability, alerting, and incident management.
  • Stay technical and partner across Nscale
    • Partner with Principal and Staff engineers on architecture and design decisions that balance automation complexity, reliability, and maintainability.
    • Stay hands-on where it counts: review designs and code, dig into incidents, and write code where it unblocks the team.
    • Collaborate with Product, Infrastructure, Platform, SRE, and UI/UX to capture requirements early, align on interfaces, and ship integrations that meet operator needs.
    • Champion AI tools like Claude and Cursor across the team as a fundamental multiplier of what your engineers can build.

Requirements

  • 8+ years of software engineering experience building and operating production systems, including 2+ years directly managing software engineers.
  • Strong technical foundation in Python and distributed systems: credible in design and code reviews, and able to guide trade-offs in infrastructure automation or workflow tooling.
  • Track record of delivering complex projects from ambiguous requirements to production, with hands-on day-2 operations experience (monitoring, incident response, performance optimization).
  • Proven people leadership: hiring, coaching, performance management, and developing engineers toward senior and staff levels.
  • Driven by building distributed systems at scale, infrastructure reliability, scalability, security, and continuous improvement.
  • Use AI tools like Claude or Cursor as a core part of your workflow and know how to raise a whole team's leverage with them.
  • Stay effective while context-switching between technical depth, delivery judgment calls, and people leadership—reviewing a design, unblocking an engineer, and running a hiring debrief in the same morning.
  • Excellent communication skills to build consensus with stakeholders, both internally and externally.

Qualifications

Strong candidates will have:

  • Experience leading teams that build workflow orchestration systems (Temporal, Airflow, Prefect, or similar).
  • Hands-on experience with infrastructure tooling: DCIMs, NetBox, OpenStack, or ERP systems.
  • Bare-metal provisioning and automation: MAAS, Ironic, IPMI, PXE boot, or network automation.
  • Experience building or managing hardware lifecycle automation: provisioning, validation, testing, or remediation workflows.
  • GPU infrastructure experience: health monitoring, burn-in testing, or cluster management.
  • HPC and networking: datacenter topology, high-performance interconnects (InfiniBand, RoCE).
  • Working knowledge of Kubernetes, Infrastructure as Code (Terraform, Pulumi), AWS, and GCP.
  • Experience scaling a team through rapid growth: hiring pipelines, onboarding, and team processes that hold up under pressure.

Benefits

  • Highly competitive US compensation package (base + bonus + equity), with performance reviews every 12 months.
  • Join one of the fastest-growing AI infrastructure companies—your work will directly shape how global AI capacity is planned and deployed.
  • Dynamic progression plan tailored to your ambitions, with support for leading critical cross-functional initiatives.
  • Human-first flexibility: trust and autonomy to shape your day around life's moments.

Pay

Salary Range: $300,000—$350,000 USD

Similar jobs