Staff Software Engineer, Resiliency (Federal)
Okta · San Francisco, CA · 4 days ago
HybridEngineering$194k/yrFull-time
Job Duties and Responsibilities
- Partner with engineering teams to design, develop, and deliver cloud-based infrastructure projects that solve complex reliability issues and help Okta’s service scale.
- Design and implement robust backend frameworks, tools, microservices, and library components for the entire engineering organization to use at scale.
- Champion innovative, scalable, and resilient solutions across the engineering organization, advising other teams on fault-tolerant system design.
- Actively explore and integrate modern AI-driven developer tooling, automated diagnostics, and predictive monitoring capabilities to streamline performance testing, accelerate debugging, and optimize service reliability.
- Conduct rigorous design and code reviews. Raise the engineering bar by implementing high-quality unit and functional tests to ensure robust programming standards.
- Participate in and lead incident Root Cause Analysis (RCA) processes to continuously improve reliability, identify systemic weaknesses, and build long-term mitigation strategies.
- Collaborate closely with Architects, QA, Product Owners, Engineering Services, and Tech Ops to deliver highly secure and dependable software.
- Help with mentoring new engineering hires and interns.
Minimum Required Knowledge, Skills, and Abilities
- 8+ years of experience as a software developer with deep, expert-level knowledge of Java (Golang experience is a strong plus).
- Demonstrated history of architecting, implementing, tuning, and debugging global cloud software. Deep understanding of microservices, cloud infrastructure (AWS preferred, Azure, or GCP), and modern container ecosystems (Docker, Kubernetes).
- Practical experience or strong interest in leveraging generative AI utilities and developer assistance tools (e.g., AI coding assistants, automated code-generation, intelligent log analytics, or predictive profiling tools) to drive development productivity and system optimization.
- Outstanding communication and leadership skills with a proven ability to coordinate high-impact technical projects, manage stakeholder alignment, and influence technical direction.
- Previous experience in a dedicated platform reliability, resilience, or core infrastructure role is highly valued.
Education and Training
- B.S. or M.S. in Computer Science, a related field, or equivalent industry experience.
Pay
The annual base salary range for this position for candidates located in the San Francisco Bay area is between: $194,000 USD - $243,000 USD.
Schedule
Not specified.
Benefits
Includes health, dental and vision insurance, 401(k), flexible spending account, and paid leave (including PTO and parental leave) in accordance with our applicable plans and policies.
Contact Information
To learn more about our Total Rewards program please visit: https://rewards.okta.com/us.