Senior Software Engineer, Platform Infrastructure
About the role
We are seeking a talented and experienced DevOps/SRE (Site Reliability Engineering) Senior Software Engineer to join our dynamic team. The ideal candidate will have a strong background in DevOps practices, cloud infrastructure management, automation, and team leadership skills.
If you have a consistent track record of architecting and building large-scale systems; enjoy solving intriguing system challenges at internet-scale; if you are innovative at heart; and have a great balance of skills in learning, organizing, building, and enjoy making an impact, this role might be a great fit for you!
Compensation
For California Only - The estimated annual salary for this position is between $280,000 - $380,000 annually. Compensation packages are based on factors unique to each candidate, including but not limited to skill set, certifications, and specific geographical location.
Benefits
- This role is eligible for health insurance
- This role is eligible for equity awards
- This role is eligible for life insurance
- This role is eligible for disability benefits
- This role is eligible for parental leave
- This role is eligible for wellness benefits
- This role is eligible for paid time off
What you’ll be doing
- You will design, implement and maintain an active-active multi-cloud infrastructure on AWS and GCP supporting business-critical systems ensuring high availability and performance with automation, delivering systems that stays reliable and performs under stress
- You will collaborate with your peers through code-reviews, ensuring best practices and aligning on technical standards to deliver consistent, high-quality solutions
- You will collaborate with security teams to ensure the integrity and security of infrastructure and applications including implementing security best practices and compliance standards
- You will manage individual project priorities, deadlines, and deliverables by leveraging agile methodologies and maintaining clear communication channels
- You will lead incident response efforts by triaging issues effectively, collaborating closely with cross-functional teams to resolve them promptly and minimize downtime
- You will implement effective incident management processes and post-incident reviews
- You will identify performance bottlenecks through detailed monitoring and profiling and optimize system resources by fine tuning configurations, scaling infrastructure and addressing latency issues
- You will drive continuous improvement initiatives by automating repetitive tasks, refining workflows and proactively addressing technical debt within the team, while driving enhancements across the organization
- You will maintain comprehensive documentation of systems, processes, and procedures while fostering a culture of knowledge sharing and contribute to the collective learning of the team
- You will participate in 24x7 on-call rotation, and be available to work with global teams in the event of critical outages