Software Engineer, Fleet Automation
NorthMark Compute & Cloud · Dallas, TX · 2 wk ago
EngineeringFull-time
About the role
NorthMark Compute & Cloud (NMC²) is seeking a Software Engineer to join the Fleet Automation team within the HPC & Infrastructure organization. This team owns the systems and tooling that keep hundreds of high-performance GPU compute nodes provisioned, configured, and operating at peak efficiency — spanning bare-metal provisioning, lifecycle management, and automated remediation at scale.
Responsibilities
- Design, build, and maintain fleet automation services and internal platforms for provisioning, configuration, and lifecycle management of large-scale GPU and CPU compute nodes.
- Develop APIs and service integrations that enable Infrastructure and Operations teams to deploy, image, validate, and decommission hardware with minimal manual intervention.
- Build and maintain backend services in Go, C#, and TypeScript with a strong focus on reliability, testability, and long-term maintainability.
- Design and evolve data models and persistent state for automation workflows, working across relational and NoSQL databases as appropriate.
- Build and maintain CI/CD pipelines that gate configuration changes, run automated hardware validation tests, and promote changes safely across environments.
- Instrument systems for observability — designing metrics, alerts, and dashboards in Prometheus and Grafana that provide real-time fleet health visibility to on-call teams.
- Participate in on-call rotations; own incident response, post-mortems, and follow-through on reliability improvements across the fleet.
- Identify systemic gaps in fleet reliability and efficiency and champion engineering solutions that reduce operational toil at scale.
Requirements
- Bachelor's Degree in Computer Science, Software Engineering, or equivalent practical experience.
- 5+ years of software engineering experience building production backend services or infrastructure automation tooling.
- Proficiency in Go, C#, or TypeScript.
- Experience designing and working with relational and NoSQL databases to support stateful automation workflows and internal platform services.
- Solid understanding of Linux systems — networking, storage, process management, and debugging on Ubuntu or RHEL variants.
- Experience building and maintaining CI/CD pipelines and observability stacks (Prometheus, Grafana, Alertmanager, ELK) in a production environment.
- Familiarity with GPU compute infrastructure and NVIDIA tooling (DCGM, nvidia-smi, NVIDIA Container Toolkit) is a strong plus.
- Exposure to event-driven architectures or messaging platforms (e.g. Kafka) is a plus for teams building automation workflows across distributed services.
- Strong communication skills and a collaborative mindset — comfortable navigating ambiguity, taking initiative, and working across Infrastructure, Operations, and Research teams.