AI Hardware Systems Manager, Annapurna Labs, Trainium Machine Learning Fleet Operations
Amazon Web Services (AWS) · Austin, TX · 3 wk ago
ConsultingFull-time
Key job responsibilities
- Build, hire, mentor, and grow a team of platform development engineers responsible for ML fleet operations across multiple accelerator platforms
- Define team roadmap and technical strategy for fleet health, automation, and data infrastructure — balancing near-term operational demands against long-term engineering investments
- Drive operational excellence by establishing metrics, SLAs, and processes that maximize platform sellability and customer experience
- Partner with hardware engineering, software engineering, and product teams to prioritize debug efforts and translate fleet learnings into permanent design fixes
- Own escalation paths for critical fleet incidents and lead cross-functional war rooms to resolution
- Influence org-level priorities by surfacing fleet-wide patterns and advocating for systemic improvements across the ML hardware portfolio
- Raise the bar on team software practices — ensuring automation is maintainable, tested, documented, and reusable at scale
- Represent fleet operations in executive reviews, providing data-driven narratives on platform health and roadmap
About the team
The MLA Fleet Operations team was formed to maintain an exceptionally high quality bar for our fleet of advanced machine learning accelerators and server products. We perfect the customer experience by developing scalable software for rapid incident response times and data visualization as well as diving deep into hardware issues as they arise.
Basic Qualifications
- Bachelor's degree in computer science, electrical engineering, or related field
- 2+ years of engineering team management experience
- Knowledge of and proficiency in the use of Python scripting language
- Experience with general troubleshooting/debugging of hardware
- Experience designing, building, operating, and managing large-scale distributed systems or web services
- 7+ years of experience in systems engineering, platform engineering, SRE, or hardware operations
Preferred Qualifications
- Experience in automating, deploying, and supporting large-scale infrastructure
- Experience in server technologies such as, thermal, mechanical, power, and signal integrity
- Experience working cross-functionally across several teams both technical and non-technical
- Experience with GPU, ML accelerator, or high-performance computing hardware
- Experience managing teams through ambiguity on new or unreleased products