Jobs · Information Technology · Washington

GPU/CPU Systems Engineer

Oracle · Seattle, WA · 3 wk ago
Information Technology$135k–$306k/yrInternship

About the Role

Oracle hardware platform development engineering is seeking a highly driven GPU/CPU Platform System Engineer at the Principal Engineer level. You will work within a small team of talented engineers leading the development and day-to-day engineering efforts for Oracle’s rapidly growing and successful Cloud AI platforms. The team has delivered the first and second generation of Oracle Cloud dedicated compute and AI platforms and is building the next generation of Cloud and Enterprise systems with record-breaking performance, security, and world-class quality using the latest merchant silicon and technologies.

In this role, you will participate in platform definition, development oversight, design reviews, system integration, performance testing, and characterization. You will collaborate closely with third-party GPU IC suppliers, internal hardware and software teams, and other stakeholders to drive Oracle’s AI Cloud platform solutions. This is a high-impact position where your work will influence OCI’s long-term architectural direction and shape the future of cloud infrastructure.

Responsibilities

  • Review and assess third-party merchant silicon used for AI Accelerator Modules.
  • Evaluate system architecture and proposed implementation path analysis; participate in platform definition and analysis.
  • Provide platform development oversight for partners and work with in-house engineering experts on design and reviews.
  • Support and guide system integration, performance testing, and characterization.
  • Collaborate with third-party GPU IC suppliers, internal hardware, software, quality assurance, cloud orchestration, security, and manufacturing teams.
  • Document and specify design intent and details in collaboration with engineering teams.
  • Participate in hardware platform security evaluations.
  • Guide internal Oracle teams on support needed to scale, monitor, and deploy products to the Cloud.
  • Assist Oracle Cloud and Support teams in root-cause analysis of hardware or software bugs through lab replication, remote debug, and collaboration with relevant teams.
  • Work with Oracle manufacturing teams to ensure hardware is secure, robustly evaluated, and performing at peak capabilities for deployment.
  • Develop, implement, and oversee day-to-day execution of AI platform development, including design plan reviews, schematics, board layout, test feature definition, and system validation plans.
  • Oversee system integration, test, and qualification; define software diagnostics features and utilize third-party and open-source AI platform tools.
  • Enhance system characterization and performance testing capabilities; support definition of in-service monitoring and error reporting needs.
  • Collaborate with hardware developers, system architects, firmware developers, GPU suppliers, storage, networking, and compute experts throughout product development and new product introduction.
  • Serve as the last level of engineering technical support for resolving complex deployed product issues.

Requirements

  • Technical hands-on experience with market-leading GPUs or alternate AI platforms from hardware and platform development, test, and characterization perspectives.
  • Ability to balance hardware performance priorities against power, cost, and cross-functional considerations.
  • Solid knowledge of AI/GPU or AI/CPU platform architecture and capabilities.
  • Strong understanding and experience running firmware and system diagnostics tools using BMC firmware, UEFI/BIOS, and Linux tools.
  • Skilled in scripting to customize tests; experience with GPU supplier test code and open-source AI test/characterization tools.
  • Experience with architecture, design, and implementation of modern server platforms, including x86 and ARM server architectures.
  • Experience with hardware development at the system, board, and FPGA level, including board-level tools and hierarchical schematic reviews.
  • Strong communication skills to articulate complex technical issues across engineering disciplines and for executive audiences.
  • Demonstrated experience debugging and root-causing complex issues with mixed hardware and software causes.
  • Experience with early-stage bring-up, power-on, platform firmware debugging, and prototype GPU/CPU/memory complex debugging.
  • Ability to isolate problems to their source and devise timely, robust solutions.
  • Experience with high-speed buses and interconnects (e.g., PCIe, DDR, Ethernet) used in modern compute and AI platforms.

Preferred Qualifications

  • 10+ years of experience in hardware design and bring-up.
  • Comfort with hardware debuggers and protocols such as PCIe, DDR, Ethernet, USB, and SPI.
  • Experience with platform-level security technologies.
  • Experience with power circuit design and signal integrity.

Pay

Hiring range in USD: $135,200 - $306,400 per year. May be eligible for bonus, equity, and compensation deferral.

Benefits

  • Comprehensive medical, dental, and vision insurance, including expert medical opinion.
  • Short-term and long-term disability insurance.
  • Life insurance and AD&D; supplemental life insurance (Employee/Spouse/Child).
  • Health care and dependent care Flexible Spending Accounts.
  • Pre-tax commuter and parking benefits.
  • 401(k) Savings and Investment Plan with company match.
  • Paid time off: Flexible Vacation for salaried employees (13 days annually for first three years, 18 days thereafter; accrual prorated for part-time).
  • 11 paid holidays.
  • Paid sick leave: 72 hours upon hire, refreshes annually, with a carryover cap of 112 hours.
  • Paid parental leave and adoption assistance.
  • Employee Stock Purchase Plan.
  • Financial planning and group legal services.
  • Voluntary benefits including auto, homeowner, and pet insurance.

Similar jobs

GPU Engineer

LenovoNorth Carolina, United States· 1 mo ago
Engineering$83/hrapply on lenovo.avature.net

System Engineer, GPU Server

Super Micro Computer Spain, S.L.San Jose, CA· 1 mo ago
Information Technology$90k–$110k/yrapply on jobs.supermicro.com