Jobs · Information Technology · Washington

Sr. Manager, Core Infrastructure Engineering

Oracle · Seattle, WA · 2 wk ago
Information Technology$122k–$306k/yrFull-time

About the Role

Oracle Cloud Infrastructure (OCI) is building some of the world's largest and most advanced GPU clusters to power the next generation of AI. The Strategic Customer Engineering (SCE) Core Infrastructure team—also known as the AI/ML Forward Deployed Infrastructure Engineering team—provides white-glove engineering and operational support to OCI's most strategic AI/ML infrastructure customers. As a trusted partner to our customers, we play a critical role in designing, deploying, operating, and optimizing the infrastructure that powers some of the largest and most demanding GPU and AI/ML environments in the world. Our team works closely with customers and internal engineering organizations to ensure exceptional reliability, performance, and scalability for mission-critical AI workloads.

We are seeking an experienced Core Infrastructure Engineering Leader to lead a high-performing engineering team responsible for delivering healthy infrastructure with optimal performance configuration. This leader will drive end-to-end technical customer execution including POCs, technical troubleshooting, post-sale customer management/support, build automated solutions for provisioning, configuring, and monitoring AI/ML infrastructure to streamline operations and enhance productivity, and optimize infrastructure performance.

Responsibilities

  • Lead, mentor, and develop a team of Core Infrastructure Engineers responsible for designing, implementing, and maintaining the infrastructure that supports our largest GPU/AI/ML customers.
  • Drive the design, development, testing, validation, and deployment readiness of our automated GPU Cluster deployment tool (like AWS Parallel Cluster, Azure Cycle Cloud) with Slurm and/or Oracle Kubernetes Engine (OKE) to streamline operations and enhance productivity.
  • Build collaborative relationships with OCI Services team, customer and sales team to deliver reliable, scalable infrastructures.
  • Act as a technical liaison between customers, core engineering teams, and support.
  • Work with OCI Strategic customers to grow our business in pre/post sales stages in a technical infrastructure expert role.
  • Take ownership of problems and work to identify solutions.
  • Ability to think through the solution and identify/document potential issues impacting your customers.
  • Optimize infrastructure performance by tuning parameters, optimizing resource utilization, and implementing caching and data pre-processing techniques.
  • Troubleshoot infrastructure performance, scalability, and reliability issues and implement solutions to mitigate risks and minimize downtime.
  • Document infrastructure designs, configurations, and procedures to facilitate knowledge sharing and ensure maintainability.
  • As a trusted customer advocate, help customers/partners understand best practices around advanced GPU solutions, and how to migrate their workloads to the cloud.
  • Educate customers of all sizes on the value proposition of Oracle Cloud and participate in deep architectural discussions to ensure solutions are designed for successful deployment in the cloud.
  • Planning & Execution: Oversee and guide multiple teams on managing complex projects or initiatives, monitoring timelines, deliverables, and budgets when applicable to ensure strategic objectives are met. Serve as a role model for appropriately delegating work, setting priorities, and ensuring alignment with business needs. Coach others on adjusting resources or project timelines in anticipation of business changes.
  • Collaboration & Partnership: Role model leading cross-functional collaborative efforts to ensure alignment of expectations and strategic objectives. Empower team to build and maintain partnerships with business leaders, stakeholders, and/or customers to address barriers and contribute to organizational success. Drive transparency and inclusivity by modeling actively seeking, listening to, and leveraging diverse perspectives.
  • Problem Solving: Share problem-solving strategies across teams, providing oversight on complex operational and/or technical issues, as needed. Coach teams on analyzing highly complex data and/or information to identify solutions to ambiguous issues and provide direction on identifying root causes to prevent recurrence of issues.
  • Continuous Learning: Pursue strategic learning opportunities to maintain expertise and apply best practices at the organizational level. Create opportunities for team members and leaders to build their expertise in new areas, coaching them to build innovative skills. Identify skill gap trends across the organization, and uphold a culture that places significant emphasis on sharing knowledge and pursuing learning opportunities that advance the organization. Evaluate efficiency of learning strategies and recommend adjustments as needed.
  • Continuous Improvement: Empower team to own the development and implementation of ideas that increase the efficiency and effectiveness of processes, protocols, and workflows across the department. Coach teams to gain buy-in for ideas and to seek feedback on approaches and methods for continued improvement. Prioritize and review the roadmap of improvement initiatives to ensure alignment with strategic direction and maximize return on investments.
  • Performance and Development: Serve as a role model for driving performance across teams through tailored feedback and coaching in alignment with performance management processes, guidelines, and expectations. Drive consistency in the application of talent development procedures and socialize performance expectations across the organization. Ensure that individual development goals are aligned with organizational strategic initiatives. Collaborate with HR to implement talent strategy through hiring and promotion processes.

Requirements

  • 7+ years of senior software engineering leadership or related experience.
  • Strong communication and collaboration skills, with the ability to work effectively in cross-functional teams and convey technical concepts to non-technical stakeholders.
  • Demonstrated leadership and people management skills.
  • Proven experience building and managing distributed/cloud software engineering solutions.
  • Demonstrated ability to:
    • Develop short, medium, and long-term plans to achieve strategic objectives.
    • Interact across functional areas with senior management or executives to ensure unit objectives are met.
    • Influence thinking and gain acceptance from others in sensitive situations.
  • BS or MS degree or equivalent experience relevant to functional area.
  • Experience using tools like Ansible, Terraform, Python, containerization technologies (e.g., Docker, Kubernetes) and orchestration tools.
  • Solid understanding of networking concepts, security principles, and best practices.
  • Excellent problem-solving skills, with the ability to troubleshoot complex issues and drive resolution in a fast-paced environment.
  • Strong Linux skills with hands-on experience in Oracle Linux/RHEL/CentOS, Ubuntu, and Debian distributions, including system administration, package management, shell scripting, and performance optimization.
  • Strong proficiency in at least one of the programming languages such as Python, Rust, Go, Java, or Scala.
  • Proven experience designing, implementing, and managing infrastructure for AI/ML or HPC workloads.

Pay

US: Hiring Range in USD from: $121,500 - $306,400 per year. May be eligible for bonus, equity, and compensation deferral.

Benefits

  • Medical, dental, and vision insurance, including expert medical opinion.
  • Short term disability and long term disability.
  • Life insurance and AD&D.
  • Supplemental life insurance (Employee/Spouse/Child).
  • Health care and dependent care Flexible Spending Accounts.
  • Pre-tax commuter and parking benefits.
  • 401(k) Savings and Investment Plan with company match.
  • Paid time off:
    • Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position.
    • Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation.
  • 11 paid holidays.
  • Paid sick leave: 72 hours of paid sick leave upon date of hire. Refreshes each calendar year. Unused balance will carry over each year up to a maximum cap of 112 hours.
  • Paid parental leave.
  • Adoption assistance.
  • Employee Stock Purchase Plan.
  • Financial planning and group legal.
  • Voluntary benefits including auto, homeowner and pet insurance.

Similar jobs