Jobs · Consulting · New York

Network Development Engineer , ML Network

Amazon Web Services (AWS) · New York, NY · 1 wk ago
ConsultingFull-time

Do you like to use your network engineering, software engineering and leadership skills to shape the future of network infrastructure for Machine Learning training and inference workloads? AWS Infrastructure Services owns the design, planning, delivery, and operation of all AWS global infrastructure. We support all AWS data centers and all of the servers, storage, networking, power, and cooling equipment that ensure our customers have continual access to the innovation they rely on.

About the role

You’ll join a diverse team of software, hardware, and network engineers, supply chain specialists, security experts, operations managers, and other vital roles. You’ll collaborate with people across AWS to help us deliver the highest standards for safety and security while providing seemingly infinite capacity at the lowest possible cost for our customers. The Core Networking team is looking for a Network Development Engineer to join our Network Fabric Engineering (NFE) team. As a Network Development Engineer in the ML team, you will be responsible for building, operating, and optimizing the performance of the Amazon networks that support AWS customers.

We own the solutions that allow racks to be aggregated and Data Centers to be interconnected. Core Networking's goal is to balance efficiency, performance, and reliability to allow customers access to their applications and data.

Responsibilities

  • Perform high-quality network and/or systems operations, deployments, scaling, technology refresh, sustaining engineering, best practices application, optimization, internal/external customer interactions, and produce appropriate technical documentation.
  • Create and implement changes on the network.
  • Work closely with automation teams to define tools that allow scaling at higher volumes.
  • Lead network projects in areas such as network sustaining engineering, deployment/implementation, scaling, technology refresh, best practices application, and/or optimization.
  • Sustain Amazon Web Services’ next-generation networks for significant AWS customers by providing critical network support to diagnose, mitigate impact, and resolve large-scale networking events.
  • Operate and troubleshoot network routing, interconnectivity, platforms, performance, and configurations issues; including necessary low-level application interaction, Unix systems, and new-age networking software tools.
  • Drive and support innovation from Top of Rack to the Network Border and everything in between.
  • Partner with broader Network/Infrastructure organizations in areas of network engineering, configuration standards, and tools.
  • Work closely with Network/Systems/Software Engineering teams to ensure fast, smooth roll-out of new designs and products, and assist with deployment and sustaining of networking software tools.
  • Maintain standards across the networking toolset and ensure full compliance with those standards and policies.
  • Deliver simple, sustainable, and repeatable solutions and processes.
  • Collaborate with highest-tier escalation engineers/resolvers.
  • Create and review documentation and processes regarding network implementation/deployments, recurring issues, new standard operating procedures, and knowledge transfer material.

A day in the life

  • Work in a 24x7 team on-call rotation, with the ability to drive into the workplace for critical events/needs.
  • Manage customers during problem resolution and operate efficiently under pressure.
  • Sit at the computer during scheduled work hours with appropriate breaks while maintaining a high level of alertness and attention to detail.
  • Travel to data center/network sites and Amazon/customer offices as needed.

About the team

AWS NFE (Network Fabric Engineering) owns datacenter network services that provide unconstrained connectivity to our customers, such as EC2, EBS, S3, and Amazon CDO Services. They own the networks internal to the datacenters end-to-end, including scaling and operational functions, for both traditional datacenters. They also own newer AWS offerings such as Outposts and Local Zones.

Requirements

  • 4+ years of major internet routing protocols experience.
  • 4+ years of working in a Linux/Unix environment.
  • 4+ years of experience in designing large-scale networks.
  • 4+ years of operating large-scale networks and operational excellence.

Preferred Qualifications

  • 1+ years of automation scripting using Python, Bash, Shell, and/or Perl.
  • Knowledge of Machine Learning and LLM fundamentals, including transformer architecture, training/inference lifecycles, and optimization techniques.
  • 4+ years of optimizing network performance with solid experience in the end-to-end stack (network and servers).

Benefits

Amazon offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance, and options for supplemental life plans), EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage, 401(k) matching, paid time off, and parental leave. Learn more about our benefits.

Pay

The base salary range for this position is $155,600 - $210,500 USD annually. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location.

Similar jobs