Jobs · Engineering · Washington

Engineering Director, Core Infrastructure

Oracle · Seattle, WA · 2 days ago
Engineering$170k–$355k/yrFull-time

About the role

The successful candidate will lead multiple teams in implementing strategies for the architecture and delivery of interdependent, scalable distributed systems that meet organizational and customer demands. They will orchestrate cross-group optimization for high-throughput, large-scale data processing; align stakeholders on scalability requirements; and oversee elastic designs and effective use of data plane platforms.

Responsibilities

  • System Design & Architecture – System Scalability: Implements strategies across multiple teams or groups for the architecture and design of interdependent scalable distributed systems, including the use of distributed state management tools, ensuring organizational and system demands are met.

  • Spearheads code and/or system optimization initiatives for large-scale data processing and high-throughput requirements across multiple areas, driving improvements that support hyper-scale systems.

  • Facilitates collaborations to define system scalability requirements, ensuring the defined requirements meet customer expectations.

  • Oversees the design of interdependent systems to scale with elasticity (e.g., effectively scaling both up and down).

  • Drives the effective use and implementation of data plane platforms for large-scale data operations.

  • System Design & Architecture – System Reliability Design: Provides strategic oversight for the architecture of fault-tolerant interdependent systems capable of withstanding in-service updates by overseeing implementation across teams of redundancy, replication, and automatic failover mechanisms.

  • Influences and sets direction for designing systems to effectively handle service disruptions (e.g., network partitions) by prioritizing consistency, availability, or partition tolerance.

  • Leads strategic optimization initiatives for handling network unreliability, including directing the design of load-shedding, throttling, and rate-limiting techniques.

  • Holds teams accountable for leveraging formal verification techniques to verify system designs and conduct peer reviews across teams.

  • Drives the design of systems that are durable and adhere to service level objectives (SLOs), developing standards for availability and durability of other computing services across the department.

  • System Design & Architecture – System Reliability Performance: Drives strategies for defining key performance indicators (KPIs) and telemetry to identify risks, gaps, or cyclical dependencies in running systems, ensuring alignment with organizational goals.

  • Directs the creation and customization of complex dashboards, telemetry systems, and alerting mechanisms that proactively monitor and ensure optimal system health across teams.

  • Correctness / Availability: Implements strategies to effectively determine if systems are meeting functional and correctness requirements, and encourages teams to identify improvement opportunities.

  • Provides thought leadership on processes for formally verifying complex features to ensure system design correctness.

  • Oversees the implementation of data replication and synchronization techniques, ensuring data integrity and availability across the organization.

  • Operational Troubleshooting & Incident Management: Provides strategic oversight for diagnosing, debugging, and resolving issues in active systems to support ongoing operation.

  • Directs strategies within teams to prevent interruptions, ensuring no maintenance windows are required for customers and users when resolving issues.

  • Drives alignment across teams for operational readiness protocol and standard operating procedures.

  • Provides expert guidance for complex incident response and root cause investigations.

  • Compliance & Security: Provides strategic guidance in architecting robust security measures to protect data and applications in multi-tenant environments, ensuring encryption techniques and access controls are implemented.

  • Oversees execution of remediation plans to address identified security gaps, promoting significant improvements and continuous advancement of security measures.

  • Drives documentation efforts and ensures cloud infrastructure compliance with industry standards and regulations.

  • Automation & Change Management: Provides strategic guidance across teams on developing and maintaining automation scripts and tools (e.g., Infrastructure as Code (IaC)) to manage cloud infrastructure.

  • Drives strategic alignment of change management plans for patching, updating, and rolling back applications, and oversees that system designs allow for automation of these processes.

Qualifications

  • Experience in leading the design and implementation of scalable distributed systems.

  • Proven track record in optimizing large-scale data processing and high-throughput systems.

  • Strong background in system reliability design, including redundancy, replication, and automatic failover mechanisms.

  • Experience in defining and implementing key performance indicators (KPIs) and telemetry systems.

  • Ability to drive formal verification techniques and peer reviews across teams.

  • Expertise in data replication and synchronization techniques to ensure data integrity and availability.

  • Strong problem-solving skills and ability to provide oversight on complex operational and technical issues.

  • Experience in creating and customizing complex dashboards, telemetry systems, and alerting mechanisms.

  • Knowledge of security measures and encryption techniques to protect data and applications in multi-tenant environments.

  • Experience in developing and maintaining automation scripts and tools (e.g., Infrastructure as Code (IaC)) to manage cloud infrastructure.

  • Strong communication and collaboration skills to facilitate cross-functional efforts and build partnerships with business leaders, stakeholders, and/or customers.

Similar jobs

Director, Core Engineering

The Walt Disney CompanyEmeryville, CA· 3 mo ago
Information Technology$243k–$315k/yrapply on disneycareers.com