System Development Engineer, AWS Manufacturing Infrastructure Services
About the role
Have you ever wondered what it would be like to build massively scalable systems that are used by the world's largest cloud infrastructures? Would you enjoy broad yet equally deep scope that impacts all AWS systems globally? AWS Manufacturing Infrastructure Services continues to pioneer and our team is architecting, building, and operating scalable services that are at the core of AWS infrastructure.
We own all AWS platforms, services, infrastructure, and tools that ensure the health of AWS hardware by testing every new system across all AWS manufacturing sites and ensure they are healthy when delivered to data centers. Our platform enables service owners such as EC2, EBS, S3, and other to deliver healthy servers for their service. Our team leads a large-scale service that sets the bar for Amazon and the industry in platform level services, effectively enabling the hardware at scale by designing and developing the software that manages the verification and testing of every server and rack in AWS manufacturing sites.
We set the bar high to ensure that AWS customers get the capacity they need to run their applications on a healthy hardware server. There are many ambiguous and difficult challenges in our fast-moving space. Our platform is mission critical and requires deep system and software expertise. You need to have the ability to work within a fast moving and startup-like environment in a large company. You will identify solutions, trying ideas, given space to fail and iterate to produce products that your customers love.
You will be a part of a team to build the next generation of platform level software and systems that enables us to deliver healthy hardware to AWS customers. You design and deliver technology solutions which solve difficult business problems.
Deeply technical engineers, who stay close to the customer as well as the systems architecture and design. They think about customer experience and the outcome. A person who works autonomously and dives deep in to a problem to deeply understand how things work, when to make subtle change, and when to disrupt the status quo to achieve the right results.
Key job responsibilities
- Own the reliability of manufacturing networking, systems, and platform infrastructure across AWS manufacturing partner sites globally.
- Develop infrastructure-as-code to stand up and manage distributed manufacturing services at new and existing sites.
- Build monitoring, alerting, and anomaly detection systems that identify infrastructure failures before they impact manufacturing throughput.
- Develop troubleshooting tools and runbooks that enable rapid diagnosis and resolution of site infrastructure issues.
- Consult with ODM/CM IT teams on network design and infrastructure standards, working across organizational boundaries where partners don't report to AWS.
- Drive site expansion readiness, ensuring new manufacturing sites are infrastructure-ready within target onboarding timelines.
- Implement security and operational best practices for manufacturing environments, including credential management, network segmentation, and access controls.
- Contribute to architecture decisions for manufacturing platform evolution.
- Develop and build systems that enables operating manufacturing infrastructure operation at scale.
About the team
We are EC2 Manufacturing Infrastructure Services. We exist to make sure that a working rack delivers to our customers in AWS data centers. Every server that powers EC2, every rack that runs AI/ML workloads, every piece of capacity that AWS customers depend on passes through our systems before it ships. We own Server Level Testing, Rack Level Testing, Firmware Upgrade services, code deployment into manufacturing sites, network architecture at contract manufacturers, and all infrastructure installed at those sites. When yield drops or a test pipeline breaks, we are the team that responds. Nothing reaches an AWS data center without passing through us first. If our services go down, manufacturing lines stop. If our infrastructure is unreliable, capacity doesn't ship. The blast radius of what we do is measured in billions of dollars of customer workloads. We are a small team with enormous scope. You will own real systems on day one. We don't have layers of abstraction between you and the problem. You build it, you ship it, you operate it, you improve it.
Basic qualifications
- 3+ years of non-internship professional software development experience
- 3+ years of designing or architecting (design patterns, reliability and scaling) of new and existing systems experience
- Experience programming with at least one modern language such as C++, C#, Java, Python, Golang, PowerShell, Ruby
- Experience in Linux OS and network troubleshooting, or experience in networking administration and troubleshooting and experience managing full application stacks from the OS up through custom applications
- Experience in enterprise scale infrastructure or development-based cloud programs/projects in a related industry
- Experience with automation or working with scripting languages like Python, Java, Perl, PHP, Ruby, Bash, Shell, or equivalent
- 3+ years of systems engineering, network engineering, or SRE experience
- Networking: TCP/IP, DNS, DHCP, VLANs, routing, firewalls, network segmentation
- Knowledge of systems engineering fundamentals (networking, storage, operating systems, firewalls)
Preferred qualifications
- Experience working in an Agile environment using the Scrum methodology
- Experience in data centers, infrastructure service providers, or related technology companies
- Knowledge of security technology and concepts (Authentication, Authorization, Single sign-on, Cryptography, etc.)
- Experience with AWS services in edge/hybrid environments (EC2, VPC, IAM, Outposts, CloudWatch, S3)
- Experience with on-premise infrastructure (servers, storage, campus networking, firewalls)
- Familiarity with PXE boot, IPMI/BMC, or bare-metal provisioning
- Experience with monitoring and observability platforms (CloudWatch, Prometheus/Grafana, or similar)
- Experience working with external partners or vendors on shared infrastructure
- Experience with manufacturing environments, industrial networking, or factory automation systems
- Understanding of on-premise IT environments
Pay
- Base salary range: $148,700.00 - $201,200.00 USD annually (USA, CA, Cupertino)
- Sign-on payments and restricted stock units (RSUs) included
- Final compensation determined based on factors including experience, qualifications, and location
Benefits
- Health insurance (medical, dental, vision, prescription)
- Basic Life & AD&D insurance and option for Supplemental life plans
- EAP, Mental Health Support, Medical Advice Line
- Flexible Spending Accounts
- Adoption and Surrogacy Reimbursement coverage
- 401(k) matching
- Paid time off
- Parental leave