Infrastructure Engineer (Storage)
Lightning AI · Buffalo-Niagara Falls Area · 1 wk ago
Information Technology$180k–$200k/yrFull-time
About the role
The Storage Infrastructure Engineer will focus on building and operating the storage systems that power large-scale AI/ML training, inference, and HPC workloads. They will work at the intersection of software, hardware, and operations, developing automation, improving reliability, and scaling distributed storage systems across our bare-metal infrastructure.
Responsibilities
- Operate and scale distributed storage systems, including VAST and S3-compatible object storage (e.g., Ceph)
- Improve performance, reliability, and efficiency of storage systems supporting large-scale AI/ML workloads
- Debug complex storage and data path issues across hardware and software layers
- Optimize storage performance to support high-throughput, low-latency AI training and inference workloads
- Build and maintain automation for provisioning, managing, and monitoring storage infrastructure
- Develop Python-based tools and workflows to reduce manual operational overhead
- Improve lifecycle management of storage clusters, from deployment through maintenance and scaling
- Manage and operate Linux-based systems in production, including bare-metal environments
- Partner with infrastructure and data center teams on hardware bring-up, upgrades, and issue resolution
- Support capacity planning, utilization tracking, and forecasting for storage systems
- Leverage monitoring and telemetry to diagnose issues and improve system performance and reliability
- Cross-functional collaboration with Infrastructure Engineering, Network Engineering, and Platform teams to integrate storage into the broader platform
- Contribute to design discussions around new infrastructure deployments and scaling strategies
- Help define best practices for operating storage systems in high-performance computing environments
Requirements
- 5+ years of experience in infrastructure engineering, systems engineering, or related roles
- Hands-on experience operating distributed storage systems (e.g., VAST, Ceph, or similar)
- Strong Linux systems experience in production environments
- Proficiency in Python or similar scripting/programming languages for automation
- Ability to debug complex issues across system boundaries (storage, OS, hardware, networking)
- Experience with storage networking protocols (e.g., NFS or similar)
- Experience with capacity planning, monitoring, and performance tuning
Preferred Experience
- Experience with VAST storage systems in production environments
- Experience operating S3-compatible object storage at scale
- Data center operations experience, including working with physical hardware
- Familiarity with AI/ML or HPC workloads and their storage requirements
- Experience supporting GPU-based workloads or large-scale compute clusters
What You'll Need
- Experience with high-performance or low-latency distributed systems
- Familiarity with high-performance data transfer technologies (e.g., RDMA, GPU Direct Storage)
Pay
The anticipated annual base salary range for this role is: $180,000 - $200,000 USD
Schedule
This role is based in one of our hubs (NYC, SF, Seattle, or London), with a minimum of 2 in-office days per week and occasional team and company offsites.