Data Center Operations Systems Engineer (Seattle)
Lambda · Quincy, WA · 2 days ago
On-siteInformation Technology$109k–$145k/yrFull-time
About the Role
This position requires presence in our Quincy, WA Data Center 5 days per week.
What You'll Do
- Ensure new server, storage and network infrastructure is properly racked, labeled, cabled, and configured
- Troubleshoot hardware and software issues in some of the world's most advanced GPU and Networking systems
- Document and update data center layout and network topology in DCIM software
- Work with supply chain & manufacturing teams to ensure timely deployment of systems and project plans for large-scale deployments
- Manage a parts depot inventory and track equipment through the delivery-store-stage-deploy-handoff process in each of our data centers
- Partner with HW Support teams to ensure data center hardware incidents with higher level troubleshooting challenges are resolved, reported on and solutions are disseminated to the large operations organization
- Work with RMA team to ensure faulty parts are returned and replacements are ordered
- Follow installation standards and documentation for placement, labeling, and cabling to drive consistency and discoverability across all data centers
Requirements
- Strong past experiences with critical infrastructure systems supporting data centers, such as power distribution, air flow management, environmental monitoring, capacity planning, DCIM software, structured cabling, and cable management
- Be familiar with carrier DIA circuit test and turn ups, fiber testing and troubleshooting
- Basic knowledge of cable optics and the different types of use
- Solid understanding of single and three phase power theories
- PDU balancing and why it is important
- Familiar with multiple cable media types and their uses
- Knowledge of cold isle and hot isle containment
- Solid understanding of server hardware and boot process
- Ability to structure, collaborate and iteratively improve on complex maintenance MOPs
- Working with product management, support, and other teams to align operational capabilities with company goals
- Translating business priorities into technical and operational requirements
- Supporting cross-functional projects where infrastructure plays a critical role
- Are action-oriented and willingness to train junior staff on best practices
- Are willing to travel for bring up of new data center locations as needed
Nice to Have
- Have 3+ years experience with critical infrastructure systems supporting data centers, such as power distribution, air flow management, environmental monitoring, capacity planning, DCIM software, structured cabling, and cable management
- Experience with/or knowledge of network topology and configurations and 400gb Infiniband architectures
- Experience with/or knowledge of DDP or SCM cluster storage systems
- Have 3+ years working with and reporting from a ticketing systems like JIRA and Zendesk
- Advanced experience with Linux administration
- Experience with High Performance Compute GPU systems (air or water cooled) - especially Nvidia NVL72
Pay
This is a salaried non-exempt role, eligible for overtime. The annual salary range for this position is $109,000 - $145,000.
Benefits
- Generous cash & equity compensation
- Health, dental, and vision coverage for you and your dependents
- Wellness and commuter stipends for select roles
- 401k Plan with 2% company match (USA employees)
- Flexible paid time off plan that we all actually use