Sr. Software Engineer – Cloud Infrastructure and Devops
About the role
This job delivers complete solutions spanning all phases of the Software Development Lifecycle (SDLC). It involves advising management on project-level issues, guiding junior engineers, operating with little supervision, and applying knowledge of technical best practices. In this role, you will be a key contributor to Venmo’s chaos engineering and business continuity efforts, building and operating the systems that ensure our infrastructure can withstand and recover from any failure scenario.
Responsibilities
- Makes technical decisions affecting multiple teams, crossing organizational boundaries
- Establishes conventions & processes to be followed by other employees
- Determines the utilization of company resources (people, money, assets) and affects the effectiveness of the company
- Handles multiple, multi-team initiatives simultaneously, using judgement to prioritize among more issues than can be handled individually
- Understands evolving industry capabilities & practices and can judiciously apply up-to-date information for optimal results
- Copes with technical challenges and provides guidance to junior engineers
- Communicates technical issues with non-technical audiences
- Spreads their behavior, principles, and knowledge as a means of improving technical results of other employees (via many means – modeling behavior, 1:1s, working sessions, quality documentation)
- Presents ideas to product management, to ideate solutions to business problems & goals
Qualifications
- 8+ years relevant experience and a Bachelor’s degree OR Any equivalent combination of education and experience.
Additional Responsibilities & Preferred Qualifications
- Bachelor’s in computer science or related field of study
- 5+ years’ experience in software development or a related field
- 3+ years’ experience operating distributed applications 24x7x365, as part of a Cloud Engineering, DevOps, and/or SRE team
- Extensive hands-on experience with designing, implementing, and supporting infrastructure (AWS experience preferred) to support global-scale services
- Deep hands-on experience with IaaS and PaaS solutions from AWS (or similar cloud provider)
- Hands-on programming and scripting (Python, Java, Bash, Go)
- Hands-on experience with containers and container orchestration: Docker, Kubernetes
- Strong communication skills with the ability to understand and explain technical issues to a non-technical audience
- Extensive hands-on experience with chaos engineering tools and frameworks (e.g., AWS Fault Injection Simulator, Gremlin, Chaos Monkey, Litmus)
- Hands-on experience designing and implementing disaster recovery solutions for distributed systems
- Experience developing and maintaining backup and restore tooling and strategies
- Experience planning and executing game day exercises or incident simulations
- Understanding of RTO/RPO requirements and how to design systems to meet them
Job Description: Cloud/DevOps Engineer – Business Continuity
Our Business Continuity team is responsible for the resilience and recoverability of Venmo’s critical systems. We design, build, and operate the infrastructure, tooling, and processes that protect the company from disruption; spanning disaster recovery automation, chaos engineering, incident game days, and backup and restore systems. Our work ensures that Venmo can withstand and rapidly recover from any failure scenario, from infrastructure outages to region-wide incidents. We operate at the intersection of cloud engineering, SRE, and reliability, and our work is foundational to the availability of every product and service at Venmo.
What You’ll Do
- Act as a hands-on contributor, while leading by example
- Mentor junior engineers
- Enjoy a high degree of independence in decision-making while being accountable for results of yourself and your team
- Part of a team dedicated to driving the scalability and reliability of Venmo’s AWS cloud infrastructure
- Contribute to initiatives: internal teams rarely have dedicated project managers. You will define the design, identify stakeholders, navigate risks and changes, coordinate colleagues’ work, implement solutions, and be accountable for timely and quality project delivery
- Troubleshoot incidents, identify root causes, fix and document problems, and implement preventive measures
- Lead by example, making meaningful contributions to the improvement of engineering teams’ production operations
- Develop and improve tools and automation to manage infrastructure and application configuration as code
- Enhance the quality, reliability, and stability of our infrastructure and operations
- Design, implement, and operate chaos engineering experiments to proactively identify and remediate system weaknesses
- Build and maintain disaster recovery automation, runbooks, and infrastructure across Venmo’s cloud environment
- Plan and execute chaos and incident game day exercises to validate system resilience and team preparedness
- Develop and maintain backup and restore tooling and processes for critical systems and data
- Define and track resilience metrics, SLOs, and recovery time objectives for owned systems
Who We Are
Click Here to learn more about our culture and community.
Commitment to Diversity and Inclusion
We are committed to building an equitable and inclusive global economy. And we can't do this without our most important asset—you. That's why we offer comprehensive, choice-based programs, to support all aspects of personal wellbeing—physical, emotional, and financial—delivering meaningful value where it matters most. We strive to create a flexible, balanced work culture with a holistic approach to benefits, including generous paid time off, healthcare coverage for you and your family, and resources to create financial security and support your mental health.