Principal Core Infrastructure Engineer
About the role
As a Principal Core Infrastructure Engineer (AI/ML Forward Deployed Infrastructure Engineer), you will play a critical role in designing, implementing, and maintaining the infrastructure that supports customers’ AI and machine learning initiatives. You will work closely with customer technical teams (data scientists, software engineers, and infrastructure/IT professionals) and internal cloud services teams to ensure AI/ML solutions are deployed efficiently, securely, and at scale. Your expertise will be crucial in optimizing infrastructure for performance, reliability, and cost-effectiveness. In this hands-on technical role, you’ll troubleshoot issues for proof-of-concept (POC) and production deployments.
Responsibilities
- Work with OCI Strategic customers to grow business in pre- and post-sales stages in a technical infrastructure expert role.
- Take ownership of problems and drive identification of solutions; think through solutions and document potential issues impacting customers.
- Design, deploy, and manage infrastructure components such as cloud resources, distributed computing systems, and data storage solutions to support AI/ML workflows.
- Collaborate with scientists and software/infrastructure engineers to understand infrastructure requirements for training, testing, and deploying machine learning models.
- Implement automation solutions for provisioning, configuring, and monitoring AI/ML infrastructure to streamline operations and enhance productivity.
- Optimize infrastructure performance by tuning parameters, optimizing resource utilization, and implementing caching and data pre-processing techniques.
- Ensure security and compliance standards are met throughout the AI/ML infrastructure stack, including data encryption, access control, and vulnerability management.
- Troubleshoot infrastructure performance, scalability, and reliability issues and implement solutions to mitigate risks and minimize downtime.
- Stay updated on emerging technologies and best practices in AI/ML infrastructure and evaluate their potential impact on systems and workflows.
- Document infrastructure designs, configurations, and procedures to facilitate knowledge sharing and ensure maintainability.
- Act as a trusted customer advocate to help customers/partners understand best practices around advanced GPU solutions and how to migrate workloads to the cloud.
- Educate customers of all sizes on the value proposition of Oracle Cloud and participate in deep architectural discussions to ensure solutions are designed for successful cloud deployment.
- Serve as a technical liaison between customers, core engineering teams, and support.
Qualifications
- Experience in scripting and automation using tools like Ansible, Terraform, Python, and/or Kubernetes.
- Experience with containerization technologies (e.g., Docker, Kubernetes) and orchestration tools (e.g., Slurm, PBS) for managing distributed systems.
- Solid understanding of networking concepts, security principles, and best practices.
- Excellent problem-solving skills with the ability to troubleshoot complex issues and drive resolution in a fast-paced environment.
- Strong communication and collaboration skills with the ability to work effectively in cross-functional teams and convey technical concepts to non-technical stakeholders.
- Strong documentation skills with experience documenting infrastructure designs, configurations, procedures, and troubleshooting steps to facilitate knowledge sharing, ensure maintainability, and enhance team collaboration.
- Strong Linux skills with hands-on experience in Oracle Linux/RHEL/CentOS, Ubuntu, and Debian distributions, including system administration, package management, shell scripting, and performance optimization.
Preferred Qualifications
- Strong proficiency in at least one of the programming languages such as Python, Rust, Go, Java, or Scala.
- Proven experience designing, implementing, and managing infrastructure for AI/ML or HPC workloads.
- Understanding of machine learning frameworks and libraries such as TensorFlow, PyTorch, or scikit-learn and their deployment in production environments.
- Familiarity with DevOps practices and tools for continuous integration, deployment, and monitoring (e.g., Jenkins, GitLab CI/CD, Prometheus).
- Strong experience with high-performance computing/GPU systems.
Core Responsibilities
- Strategic Thought Leadership: Demonstrated ability to think strategically about business, products, and technical challenges.
- Business Experience: Understand the challenges in working with large customers.
- Planning & Execution: Manages and coordinates moderately complex tasks, monitoring timelines and deliverables to ensure timely completion and adherence to requirements for a moderately sized project or initiative; efficiently delegates, monitors, and prioritizes work across multiple projects, providing technical oversight and adjusting plans to address shifts in resources or timelines.
- Collaboration & Partnership: Collaborates across the organization to align on expectations and achieve shared objectives; leverages understanding of business leaders, stakeholders, and/or customers to ensure proposed solutions meet their needs; supports inclusivity by actively seeking and listening to diverse perspectives.
- Problem Solving: Identifies and addresses moderately complex issues by analyzing a wide range of data and/or information to identify solutions in accordance with standard practices; proactively escalates unresolved or critical issues with a thorough assessment and suggests potential solutions; reviews, contributes to, and documents problem-solving strategies.
- Continuous Learning: Pursues learning opportunities to expand knowledge and skills and/or tools in new areas and stays abreast of the latest industry trends and best practices; proactively seeks and leverages ongoing feedback and training to improve skills; coaches and mentors junior team members, fostering continuous learning and knowledge sharing within and across teams.
- Continuous Improvement: Develops ideas, recommends updates, and/or collaborates on the implementation of process improvements to increase the efficiency and effectiveness of processes, protocols, and workflows across teams; evaluates the impact on key stakeholders; solicits feedback from others on ideas for alternative approaches and methods for continued improvement.
- Performance and Development: Contributes to the talent development pipeline by participating in candidate interviews, assessing candidates, and providing hiring recommendations.
Pay
U.S. range: $114,600 – $234,600 per year. May be eligible for bonus, equity, and compensation deferral. Oracle maintains broad salary ranges to account for variations in knowledge, skills, experience, market conditions and locations, as well as reflect differing products, industries and lines of business. Candidates are typically placed into the range based on the preceding factors as well as internal peer equity.
Benefits
- Medical, dental, and vision insurance, including expert medical opinion.
- Short-term and long-term disability insurance.
- Life insurance and accidental death & dismemberment (AD&D) insurance.
- Supplemental life insurance (Employee/Spouse/Child).
- Health care and dependent care flexible spending accounts.
- Pre-tax commuter and parking benefits.
- 401(k) Savings and Investment Plan with company match.
- Paid time off: Flexible Vacation for salaried employees; accrued vacation for others based on hours worked and tenure.
- 11 paid holidays.
- Paid sick leave: 72 hours upon hire; refreshes annually up to a 112-hour cap.
- Paid parental leave.
- Adoption assistance.
- Employee Stock Purchase Plan.
- Financial planning and group legal services.
- Voluntary benefits including auto, homeowner, and pet insurance.