Principal Site Reliability Engineer
About the role
Solve complex problems related to infrastructure cloud services and build automation to prevent problem recurrence. Design, write, and deploy software to improve the availability, scalability, and efficiency of Oracle products and services. Design and develop designs, architectures, standards, and methods for large-scale distributed systems. Facilitate service capacity planning and demand forecasting, software performance analysis, and system tuning.
You will provide cloud operations for Oracle National Security Realms. You’ll be part of a dynamic team with a broad knowledge of how Oracle’s cloud platform works. You’ll partner with customer support, service owners, and engineering teams around the globe to ensure high-quality service for customers.
This role involves working a 24/7 shift rotation with on-call duties, including nights, weekends, and public holidays.
Responsibilities
- Serve as escalation points for junior site reliability engineers during complex or high-impact incidents.
- Manage and execute complex manual Change Management tickets, working closely with service teams to ensure safety and minimal disruption to services.
- Support the on-boarding of new services and tools, ensuring they are operationally ready and properly integrated.
- Provide mentorship and training to SREs, helping build team capability and confidence.
- Create and maintain clear, useful documentation for operational processes and system support.
- Identify areas of manual work and drive automation to reduce toil and improve efficiency.
- Automate tasks to enable continuous delivery and ensure continuous availability with minimal human overhead.
- Recognize unsafe or inefficient practices and work with teams to design safer, more effective solutions.
- Complete change requests to enable new functionality and maintain realm compliance.
- Ensure timely resolution of incidents, service requests, and change requests.
- Collaborate with global service and engineering teams.
- Define and drive change management, continuous integration, and deployment best practices.
- Help create and maintain real-world production architectures, scalability, and system design.
- Use a methodical approach to troubleshoot large, complex, interconnected systems.
- Specific experience working with deployment of AI infrastructure to include clustered GPUs, LLM deployment and maintenance, and understanding of model integration for customer solutions.
Requirements
- US Citizenship and possession and maintenance of TS/SCI w/Poly security clearance.
- Reside in Austin, TX or Reston, VA.
- Technology-related bachelor’s degree and/or equivalent work experience.
- A desire to learn and keep up with modern technologies.
- Proficient with writing services/task automation in any modern development language (e.g., Python, Bash, Ruby, Perl, JavaScript, or Java).
- Familiarity with core protocols and OSI model (DNS, DHCP, HTTP, TCP/IP).
- Deep knowledge of Linux or Unix OS internals and host-based networking.
- Familiarity with configuration management solutions such as Chef, Puppet, etc.
- Experience with devising, managing, and extending monitoring solutions for large-scale environments.
- Knowledge of cloud computing concepts.
- Experience working in a mission-critical environment (Operations, Technical Support, NOC, etc.).
- Proficient communication skills (writing, organization, learning exchange).
- Experience executing tasks under change management procedures.
- Experience resolving auto-cut and manual alarms following runbooks.
- A focus on customer satisfaction.
Preferred Qualifications
- Certifications in VMware or other hypervisors.
- Certifications in Cisco (e.g., CCNA, CCNP).
- Certifications in CISSP or other security-related areas.
- Certifications in Oracle Databases or RAC.
Pay
Hiring range in USD from: $84,900 - $209,500 per year. May be eligible for bonus and equity.
Benefits
- Medical, dental, and vision insurance, including expert medical opinion.
- Short-term disability and long-term disability.
- Life insurance and AD&D.
- Supplemental life insurance (Employee/Spouse/Child).
- Health care and dependent care Flexible Spending Accounts.
- Pre-tax commuter and parking benefits.
- 401(k) Savings and Investment Plan with company match.
- Paid time off: Flexible Vacation for salaried employees; accrued vacation for others (13 days annually for the first three years, 18 days annually thereafter; prorated for part-time).
- 11 paid holidays.
- Paid sick leave: 72 hours upon hire, refreshes annually, carries over up to 112 hours.
- Paid parental leave.
- Adoption assistance.
- Employee Stock Purchase Plan.
- Financial planning and group legal.
- Voluntary benefits including auto, homeowner, and pet insurance.
Schedule
24/7 shift rotation with on-call duties, including nights, weekends, and public holidays.