Senior AI Solution Architect
About the role
Intel is a company of bold and curious inventors and problem solvers who create some of the most astounding technology advancements and experiences in the world. We focus on designing, implementing, and optimizing AI accelerator systems and cloud infrastructure for large-scale machine learning workloads.
Responsibilities
- Design and optimize AI accelerator systems (Gaudi, GPU clusters) for production ML workloads.
- Debug complex PCIe, memory subsystem, and interconnect issues in AI clusters.
- Validate and integrate cutting-edge GPUs and AI accelerator platforms.
- Lead platform bring-up and validation for next-generation AI hardware.
- Develop comprehensive test plans for AI systems.
- Collaborate with OEM vendors on BMC firmware integration and system stability.
- Perform full-stack debugging across hardware, firmware, and software layers.
- Develop automated testing frameworks and monitoring solutions.
- Create diagnostic tools and APIs for system health monitoring.
- Mentor junior engineers and data center technicians.
- Lead cross-functional teams through complex technical challenges.
- Coordinate with hardware, firmware, and software teams on platform readiness.
- Drive technical decisions and architectural improvements.
Requirements
- Bachelor’s and 6+ years, or Master’s and 4+ years, or PhD and 2+ years in Computer Science, Electrical Engineering, or related field.
- 5+ years of experience in system engineering, platform validation, or related roles.
- 3+ years of experience successfully bringing up and debugging high-performance AI clusters.
- 3+ years of experience resolving complex system-level issues in production AI/ML environments.
- 3+ years of AI cluster design, validation, and production deployment experience.
Preferred Qualifications
- Experience with Intel platforms (Xeon, Gaudi) or similar GPU or AI accelerators.
- Familiarity with cloud deployment and containerization.
- Expert-level Python programming.
- Experience with AI/ML frameworks: vLLM, PyTorch, TensorFlow, OpenMPI.
- Proficiency with system tools: Linux/Unix administration, Docker, shell scripting.
- Deep understanding of PCIe, memory subsystems, AI accelerators.
- Knowledge of protocols: Redfish, IPMI, BMC management.
- Computer architecture and microprocessor design.
- AI/ML workload optimization and deployment.
- System-level debugging and validation methodologies.
- Enterprise platform security and manageability.
Benefits
We offer a total compensation package that ranks among the best in the industry. It consists of competitive pay, stock bonuses, and benefit programs which include health, retirement, and vacation.
Pay
Annual Salary Range for jobs which could be performed in the US: $170,500.00 - 315,490.00 USD. The range reflects the minimum and maximum target compensation for the position across all US locations. Within the range, individual pay is determined by work location and additional factors, including job-related skills, experience, and relevant education or training.
Schedule
This role will require an on-site presence (Shift 1). Primary location: US, Oregon, Hillsboro.