Systems Development Engineer II, AWS Managed Operations (MO) AWSOM Team
About the role
AWS Managed Operations (MO) is building a best-in-class engineering and operations team to own the day-to-day operations for AWS Regions, improving availability, reliability, latency, performance, and efficiency. We are looking for Systems Development Engineers who can balance day-to-day operations with long-term software engineering to reduce operational toil, and who enjoy constantly learning across the wide range of systems that make up one of the world's largest cloud providers.
Key job responsibilities
- Collaborate across diverse teams, projects, and environments to have a firsthand impact on our global customer base
- Build solutions to innovate our approaches to agentic-first interfaces
- Solve challenging technical problems, often ones not solved before, at every layer of the stack
- Demonstrate ability to leverage Generative AI tools and AI-assisted development environments to accelerate prototyping, development, validation, and testing workflows
- Use AI-powered code generation and review tools to rapidly iterate on infrastructure automation scripts, configuration templates, and systems tooling, and to automate investigative, diagnostic, or operational tasks
- Apply GenAI capabilities to accelerate test case generation, integration testing, and validation of complex infrastructure components — reducing cycle times without sacrificing quality or security rigor
- Use GenAI tools to rapidly synthesize technical documentation, compliance frameworks, and architectural patterns — accelerating research and decision-making in ambiguous problem spaces
- Maintain awareness of responsible GenAI use in security-sensitive environments, including data handling boundaries, model limitations, and appropriate human-in-the-loop validation practices
A day in the life
You will split your time approximately 50/50 between operating production systems and driving long-term improvements to reliability, availability, and performance.
A typical week includes:
- Root cause analysis & remediation — Investigate production issues such as deployment failures, identify underlying bugs, and implement fixes
- Systems-level problem solving — Identify patterns across incidents, design solutions for entire classes of problems, and collaborate with your team to refine designs
- SLO stewardship — Evaluate Service Level Objective effectiveness, partner with stakeholders to validate thresholds, and update infrastructure as code to keep monitoring meaningful and actionable
- Systems automation for operational excellence — Design and build systems that improve fleet management, such as safely migrating workloads to more optimal hardware types, delivering measurable performance gains
About the team
Our team has a broad mix of experience levels and tenures, and we celebrate knowledge-sharing and mentorship. Senior members provide one-on-one mentoring and thorough, kind code reviews. We care about your career growth and assign projects that help develop engineering expertise, empowering you to take on more complex tasks in the future.
AWS values diverse experiences. Even if you do not meet all preferred qualifications, we encourage you to apply. If your career is just starting, hasn't followed a traditional path, or includes alternative experiences, don't let it stop you.
AWS Infrastructure Services owns the design, planning, delivery, and operation of all AWS global infrastructure — we keep the cloud running. We support all AWS data centers and the servers, storage, networking, power, and cooling equipment that ensure customers have continual access to innovation.
We value work-life harmony and strive for flexibility as part of our working culture. We offer endless knowledge-sharing, mentorship, and other career-advancing resources to help you develop into a better-rounded professional.
Basic qualifications
- 3+ years of designing or architecting (design patterns, reliability and scaling) of new and existing systems experience
- 3+ years of administrative experience in networking, storage systems, operating systems and hands-on systems engineering experience
- Experience programming with at least one modern language such as Python, Ruby, Golang, Java, C++, C#, Rust
- Experience developing, deploying and managing AI products at scale
Preferred qualifications
- 5+ years of administrative experience in networking, storage systems, operating systems and hands-on systems engineering experience
- 2+ years of non-internship professional software development experience
- Experience in networking, storage systems, operating systems and hands-on systems engineering
- Experience with distributed systems at scale
Pay
Base salary range: $129,200.00 - $174,800.00 USD annually (USA, VA, Herndon). Total compensation package includes sign-on payments and restricted stock units (RSUs). Final compensation determined by experience, qualifications, and location.
Benefits
- Health insurance: medical, dental, vision, prescription
- Basic Life & AD&D insurance with option for Supplemental life plans
- Employee Assistance Program (EAP) and Mental Health Support
- Medical Advice Line
- Flexible Spending Accounts
- Adoption and Surrogacy Reimbursement coverage
- 401(k) matching
- Paid time off
- Parental leave