Staff App Ops Engineer
Intuit · Mountain View, CA · Yesterday
On-siteEngineering$203k–$274k/yrFull-time
Responsibilities
- Own "always-on" reliability at scale — drive operational excellence for distributed systems operating on AWS Cloud serving millions of small business customers with 5-9’s of availability.
- Architect cloud-native data systems — design highly available, secure, performant infrastructure at multi-million-user scale.
- Build the future of AIOps — design and ship AI-powered tools that auto-detect, triage, and remediate issues, multiplying your team's impact.
- Lead incident response and Production Support — Part of on-call rotation to lead Production Operations in engaging incidents, and resolve production issues
- Resiliency practices— Champion and innovate resiliency best practices to validate FMEA, chaos engineering practices and disaster recovery runbooks.
- Observability — Elevate the Operational Excellence with best-in-class monitoring and resolution techniques to reduce MTTD/MTTR.
- Automate relentlessly —Reduce developer toil with automating manual steps in the standard operating procedures
- Drive complex, cross-team initiatives — sequence dependencies, manage risk, and deliver ambitious programs end to end.
- Influence beyond your team — partner with platform and leadership to remove organizational blockers and accelerate paved-road adoption (Intuit wide standards)
- Shape engineering culture — Take active participation in RCA’s, contributing to best practices, and do not hesitate to debate on technology choices.
Qualifications
- 8+ years of hands-on development and operational experience building and maintaining infrastructure in AWS, with deep expertise in at least one SRE discipline (automation, monitoring, DevOps, or cloud operations).
- Deep AWS and Kubernetes expertise, including hosting on cloud, high-availability architectures, DR strategies, and hands-on knowledge of Docker, Kubernetes, and ArgoCD.
- Proficiency in a modern programming/scripting language — Go, Java, Python, or Ruby — for both application development and DevOps automation.
- Extensive observability and performance experience — monitoring, troubleshooting, and tuning using tools like Splunk, Wavefront, AppDynamics, Prometheus, or distributed tracing.
- Strong grasp of SSDLC and CI/CD pipelines, with the judgment to build and evolve them safely at scale.
- Proven technical leadership with data-backed evidence of impact and a track record of driving initiatives independently.
- Bias for action — takes initiative, unblocks themselves, and delivers work incrementally to gather feedback and iterate.
- Sharp production diagnostic skills — a passion for and demonstrated ability to resolve both pre-production and production issues under pressure.
- A strong collaborator — communicates clearly, takes feedback constructively, and is willing to do unglamorous work when it's what the team needs.