Sr. Production Engineer
Yahoo · United States · Yesterday
RemoteRemoteManagement$128k–$267k/yrFull-time
About the role
The Yahoo commerce team is seeking experienced DevOps/Cloud Infrastructure engineers with expertise in AWS and a strong knowledge of Web applications.
Responsibilities
- Analyze current frameworks and infrastructures to oversee high-capacity distributed systems, pinpoint potential enhancements, and establish architectural operational standards.
- Engage with development teams to shape the product trajectory, offering senior-level expertise in scaling, capacity management, security, and operational monitoring.
- Address intricate networking and system difficulties to bolster overall platform stability and performance.
- Leverage AI-assisted coding tools (e.g., Amazon Q) to accelerate the design, implementation, and refinement of Infrastructure as Code (Terraform) scripts and custom system automation tools.
- Execute frontend visibility strategies, including Real User Monitoring (RUM), performance tracking for Core Web Vitals, and managing error budgets using Datadog or CloudWatch.
- Coordinate with frontend developers to establish performance thresholds, deployment protocols, and tools that enhance the developer experience.
- Drive progressive delivery—feature flags, canary releases, and A/B deployment strategies.
- Establish automated, AI-driven diagnostic pipelines using cloud tools (such as AWS Bedrock / Anthropic Claude) to ingest log streams, accelerate root-cause analysis, and streamline incident post-mortems.
- Take a leading role in the on-call cycle to troubleshoot and resolve live production issues, driving follow-up efforts to prevent recurrence.
- Mentor junior engineers within the immediate squad, leading small team initiatives and running knowledge-sharing sessions on cloud engineering and modern DevOps standards.
Qualifications
- A minimum of five years of professional experience across SRE, Production Engineering, DevOps, and software development (or equivalent demonstrated engineering capacity).
- Hands-on experience utilizing AI pair-programming assistants (e.g., Amazon Q, GitHub Copilot, or Cursor) to optimize development cycles, review configurations, and accelerate script generation.
- Proficiency in developing with at least one language such as Go, Python, Java, or Node.js.
- Hands-on experience with AWS cloud platforms, Docker & container deployment (ECS or EKS) for at least 2 years.
- Solid grasp of networking protocols and security standards, including TLS/SSL, HSTS, HTTPS enforcement, and certificate management.
- Proven track record designing and maintaining reusable Infrastructure as Code (IaC) solutions, specifically Terraform.
- Experience utilizing CI/CD pipelines, particularly GitHub Actions.
- Experience utilizing Datadog or CloudWatch to maintain system observability.
Preferred Qualifications
- Extensive background in architecting and overseeing significant AWS infrastructure, including ECS, EKS, and multi-region/zone deployments.
- In-depth expertise in UNIX/Linux internals along with various troubleshooting utilities for networking and application stack analysis.
- Practical experience leveraging GitHub Actions for automated workflows.
- Knowledgeable in deploying software load balancing solutions such as Nginx.
- Adept with reliability and monitoring tools including Prometheus, Splunk, OpenTelemetry.
- Proficient in managing storage and database solutions like Redis, etc.