Jobs · Information Technology · Tennessee

Director, Core Infrastructure Engineering

Oracle · Nashville, TN · 1 mo ago
Information Technology$170k–$355k/yrFull-time

About the role

The successful candidate will lead multiple teams in implementing strategies for the architecture and delivery of interdependent, scalable distributed systems that meet organizational and customer demands. They will orchestrate cross-group optimization for high-throughput, large-scale data processing; align stakeholders on scalability requirements; and oversee elastic designs and effective use of data plane platforms.

Responsibilities

  • System Design & Architecture – System Scalability: Implements strategies across multiple teams or groups for the architecture and design of interdependent scalable distributed systems, including the use of distributed state management tools, ensuring organizational and system demands are met.

  • Spearheads code and/or system optimization initiatives for large-scale data processing and high-throughput requirements across multiple areas, driving improvements that support hyper-scale systems.

  • Facilitates collaborations to define system scalability requirements, ensuring the defined requirements meet customer expectations.

  • Oversees the design of interdependent systems to scale with elasticity (e.g., effectively scaling both up and down).

  • Drives the effective use and implementation of data plane platforms for large-scale data operations.

  • System Design & Architecture – System Reliability Design: Provides strategic oversight for the architecture of fault-tolerant interdependent systems capable of withstanding in-service updates by overseeing implementation across teams of redundancy, replication, and automatic failover mechanisms.

  • Influences and sets direction for designing systems to effectively handle service disruptions (e.g., network partitions) by prioritizing consistency, availability, or partition tolerance.

  • Leads strategic optimization initiatives for handling network unreliability, including directing the design of load-shedding, throttling, and rate-limiting techniques.

  • Holds teams accountable for leveraging formal verification techniques to verify system designs and conduct peer reviews across teams.

  • Drives the design of systems that are durable and adhere to service level objectives (SLOs), developing standards for availability and durability of other computing services across the department.

  • System Design & Architecture – System Reliability Performance: Drives strategies for defining key performance indicators (KPIs) and telemetry to identify risks, gaps, or cyclical dependencies in running systems, ensuring alignment with organizational goals.

  • Directs the creation and customization of complex dashboards, telemetry systems, and alerting mechanisms that proactively monitor and ensure optimal system health across teams.

  • Correctness / Availability: Implements strategies to effectively determine if systems are meeting functional and correctness requirements, and encourages teams to identify improvement opportunities.

  • Provides thought leadership on processes for formally verifying complex features to ensure system design correctness.

  • Oversees the implementation of data replication and synchronization techniques, ensuring data integrity and availability across the organization.

  • Operational Troubleshooting & Incident Management: Provides strategic oversight for diagnosing, debugging, and resolving issues in active systems to support ongoing operation.

  • Directs strategies within teams to prevent interruptions, ensuring no maintenance windows are required for customers and users when resolving issues.

  • Drives alignment across teams for operational readiness protocol and standard operating procedures.

  • Provides expert guidance for complex incident response and root cause investigations.

  • Compliance & Security: Provides strategic guidance in architecting robust security measures to protect data and applications in multi-tenant environments, ensuring encryption techniques and access controls are implemented.

  • Oversees execution of remediation plans to address identified security gaps, promoting significant improvements and continuous advancement of security measures.

  • Drives documentation efforts and ensures cloud infrastructure compliance with industry standards and regulations.

  • Automation & Change Management: Provides strategic guidance across teams on developing and maintaining automation scripts and tools (e.g., Infrastructure as Code (IaC)) to manage cloud infrastructure.

  • Drives strategic alignment of change management plans for patching, updating, and rolling back applications, and oversees that system designs allow for automation of these processes.

Qualifications

  • Experience in leading the design and implementation of scalable distributed systems.

  • Strong understanding of system reliability and resilience principles.

  • Expertise in data plane platforms and large-scale data processing.

  • Ability to drive strategic initiatives and set direction for system design.

  • Proven track record in managing and optimizing system performance and reliability.

  • Experience in creating and customizing complex dashboards and alerting mechanisms.

  • Knowledge of key performance indicators (KPIs) and telemetry systems.

  • Experience in security architecture and compliance.

  • Strong leadership and collaboration skills.

  • Ability to drive continuous improvement and innovation.

Similar jobs

Director, Core Engineering

The Walt Disney CompanyEmeryville, CA· 3 mo ago
Information Technology$243k–$315k/yrapply on disneycareers.com