Principal Core Infrastructure Engineer- Nashville,TN
About the role
This role is based in Nashville, TN. The Oracle Cloud Infrastructure (OCI) team builds and manages a suite of massive scale, integrated cloud services in a broadly distributed, multi-tenant cloud environment. OCI is committed to providing the best in cloud products that meet the needs of customers tackling some of the world’s biggest challenges. This is an opportunity to join the Virtual Networking Control Plane organization, a core OCI team that owns and manages networking resources across virtual and functional layers. The team serves internal OCI customers and external customers using virtual networking APIs and workflows, delivering Network-as-a-Service that handles planning, provisioning, lifecycle management, and security of customers’ network infrastructure. The role helps shape the management of network infrastructure by providing the flexibility and simplicity of a virtual network with the manageability and control of a software-defined network (SDN).
Responsibilities
- System Design & Architecture – System Scalability:
- Lead development and begin to architect components of scalable, distributed systems that support horizontal and vertical scaling to meet system demands, including leveraging distributed state management tools.
- Optimize code and/or systems for large-scale data processing and high-throughput requirements to support hyper-scale systems.
- Define scalability requirements for owned components and ensure design and implementation requirements are met.
- Design systems to scale with elasticity (e.g., effectively scaling both up and down).
- Leverage data plane platforms to effectively handle large-scale data retrieval, storage, and processing.
- Design performance and load testing.
- System Design & Architecture – System Reliability Design:
- Build and design fault-tolerant components and systems capable of withstanding in-service updates by implementing redundancy, replication, and automatic failover mechanisms.
- Design systems to effectively handle service disruptions (e.g., network partitions) by prioritizing consistency, availability, or partition tolerance.
- Implement and optimize approaches to handle network unreliability, including load-shedding, throttling, and rate-limiting.
- Design components and systems that are durable and adhere to service level objectives (SLOs), setting expectations for availability and durability of other computing services within the department.
- System Design & Architecture – System Reliability Performance:
- Define key performance indicators (KPIs) and telemetry to identify gaps or issues in running systems.
- Build and customize moderately complex dashboards, telemetry systems, and alerting mechanisms to proactively monitor components and system health.
- System Design & Architecture – Correctness / Availability:
- Design and implement functional and correctness requirements for feature sets and/or systems in new or existing systems.
- Design complex test scenarios (e.g., fault-injection, brown-out) to evaluate system correctness.
- Implement data replication and synchronization techniques to maintain data integrity and availability.
- Operational Troubleshooting & Incident Management:
- Take a proactive role in diagnosing, debugging, and resolving issues in active components and systems to support ongoing operation, and mentor others in these processes.
- Implement strategies to prevent interruptions, ensuring no maintenance windows are required for customers and users when resolving issues.
- Maintain expertise in owned components and systems to ensure effective troubleshooting and performance.
- Meet operational readiness expectations through design and implementation.
- Serve in operational support rotations, providing guidance in incident response and root cause investigations.
- Compliance & Security:
- Implement robust security measures to protect data and applications in multi-tenant environments, including encryption techniques and access controls.
- Execute remediation plans to address identified security gaps.
- Ensure cloud infrastructure is in compliance with industry standards and regulations and that documentation is up to date.
- Automation & Change Management:
- Develop and maintain automation scripts and tools (e.g., Infrastructure as Code (IaC)) to manage cloud infrastructure.
- Create and adhere to change management plans for patching, updating, and rolling back applications, and begin designing systems and components to allow for automation of these processes.
- Planning & Execution:
- Manage and coordinate moderately complex tasks, monitoring timelines and deliverables to ensure timely completion and adherence to requirements for a moderately-sized project or initiative.
- Efficiently delegate, monitor, and prioritize work across multiple projects, providing technical oversight and adjusting plans to address shifts in resources or timelines.
- Collaboration & Partnership:
- Collaborate across the organization to align on expectations and achieve shared objectives.
- Leverage understanding of business leaders, stakeholders, and/or customers to ensure proposed solutions meet their needs.
- Support inclusivity by actively seeking and listening to diverse perspectives, ensuring others feel heard and respected.
- Problem Solving:
- Identify and address moderately complex issues by analyzing a wide range of data and/or information to identify solutions in accordance with standard practices.
- Proactively escalate unresolved or critical issues with a thorough assessment and suggest potential solutions.
- Review, contribute to, and document problem solving strategies.
- Continuous Learning:
- Pursue learning opportunities to expand knowledge and skills and/or tools in new areas and stay abreast of the latest industry trends and best practices.
- Proactively seek and leverage ongoing feedback and training to improve skills.
- Coach and mentor junior team members, fostering continuous learning and knowledge sharing within and across teams.
- Continuous Improvement:
- Develop ideas, recommend updates, and/or collaborate on the implementation of process improvements to increase the efficiency and effectiveness of processes, protocols, and workflows across teams, and evaluate the impact on key stakeholders.
- Solicit feedback from others on ideas for alternative approaches and methods for continued improvement.
- Performance and Development:
- Contribute to the talent development pipeline by participating in candidate interviews, assessing candidates, and providing hiring recommendations.
Qualifications
- Experience leading development and beginning to architect components of scalable, distributed systems.
- Experience optimizing code and/or systems for large-scale data processing and high-throughput requirements.
- Experience defining scalability requirements and designing systems for elasticity.
- Experience designing fault-tolerant systems with redundancy, replication, and failover.
- Experience implementing load-shedding, throttling, and rate-limiting to handle network unreliability.
- Experience defining KPIs, telemetry, dashboards, and alerts for system health monitoring.
- Experience designing correctness tests (e.g., fault-injection, brown-out) and implementing data replication and synchronization.
- Experience diagnosing and resolving production issues and mentoring peers.
- Experience implementing robust security controls and executing remediation in multi-tenant environments.
- Experience developing and maintaining automation scripts and Infrastructure as Code (IaC).
- Experience managing moderately complex projects and coordinating work across teams.
- Experience collaborating across organizations to align on objectives and meet stakeholder needs.
Pay
U.S.: Hiring range from $114,600 to $234,600 per year. May be eligible for bonus, equity, and compensation deferral. Oracle maintains broad salary ranges to account for variations in knowledge, skills, experience, market conditions, locations, and internal peer equity.
Benefits
- Medical, dental, and vision insurance, including expert medical opinion
- Short term disability and long term disability
- Life insurance and AD&D
- Supplemental life insurance (Employee/Spouse/Child)
- Health care and dependent care Flexible Spending Accounts
- Pre-tax commuter and parking benefits
- 401(k) Savings and Investment Plan with company match
- Paid time off: Flexible Vacation for salaried employees; accrued vacation for others; 11 paid holidays
- Paid sick leave: 72 hours upon hire, refreshes annually up to a maximum cap of 112 hours
- Paid parental leave
- Adoption assistance
- Employee Stock Purchase Plan
- Financial planning and group legal
- Voluntary benefits including auto, homeowner and pet insurance