Sr. Software Engineer – Cloud Infrastructure and Devops
PayPal · Chicago, IL · 1 mo ago
Engineering$131k–$194k/yrFull-time
About the role
Job Summary: This job delivers complete solutions spanning all phases of the Software Development Lifecycle (SDLC). It involves advising management on project-level issues, guiding junior engineers, operating with little supervision, and applying knowledge of technical best practices. In this role, you will be a key contributor to Venmo’s chaos engineering and business continuity efforts, building and operating the systems that ensure our infrastructure can withstand and recover from any failure scenario.
Essential Responsibilities
- Delivers complete solutions spanning all phases of the Software Development Lifecycle (SDLC) (design, implementation, testing, delivery and operations), based on definitions from more senior roles.
- Advises immediate management on project-level issues
- Guides junior engineers
- Operates with little day-to-day supervision, making technical decisions based on knowledge of internal conventions and industry best practices
- Applies knowledge of technical best practices in making decisions
Additional Responsibilities & Preferred Qualifications
- Engineering at Venmo
- Cloud/DevOps Engineer – Business Continuity
What You’ll Do
- Act as a hands-on contributor, while leading by example
- Mentor junior engineers
- Enjoy a high degree of independence in decision-making while being accountable for results of yourself and your team
- Part of a team dedicated to driving the scalability and reliability of Venmo’s AWS cloud infrastructure
- Contribute to initiatives: internal teams rarely have dedicated project managers. You will define the design, identify stakeholders, navigate risks and changes, coordinate colleagues' work, implement solutions, and be accountable for timely and quality project delivery
- Troubleshoot incidents, identify root causes, fix and document problems, and implement preventive measures
- Lead by example, making meaningful contributions to the improvement of engineering teams' production operations
- Develop and improve tools and automation to manage infrastructure and application configuration as code
- Enhance the quality, reliability, and stability of our infrastructure and operations
- Design, implement, and operate chaos engineering experiments to proactively identify and remediate system weaknesses
- Build and maintain disaster recovery automation, runbooks, and infrastructure across Venmo’s cloud environment
- Plan and execute chaos and incident game day exercises to validate system resilience and team preparedness
- Develop and maintain backup and restore tooling and processes for critical systems and data
- Define and track resilience metrics, SLOs, and recovery time objectives for owned systems
Who We're Looking For
- Bachelor’s in computer science or related field of study
- 5+ years’ experience in software development or a related field
- 3+ years’ experience operating distributed applications 24x7x365, as part of a Cloud Engineering, DevOps, and/or SRE team
- Extensive hands-on experience with designing, implementing, and supporting infrastructure (AWS experience preferred) to support global-scale services
- Deep hands-on experience with IaaS and PaaS solutions from AWS (or similar cloud provider)
- Hands-on programming and scripting (Python, Java, Bash, Go)
- Hands-on experience with containers and container orchestration: Docker, Kubernetes
- Strong communication skills with the ability to understand and explain technical issues to a non-technical audience
- Extensive hands-on experience with designing and implementing disaster recovery solutions for distributed systems
- Experience developing and maintaining backup and restore tooling and strategies
- Experience planning and executing game day exercises or incident simulations
- Understanding of RTO/RPO requirements and how to design systems to meet them