Jobs · Engineering · Texas

AI and Machine Learning Engineer

Hewlett Packard Enterprise · Spring, TX · 3 wk ago
EngineeringFull-time

This role is designed as 'Hybrid' with a requirement to work on average 2 days per week from an HPE office.

About the role

High Performance Computing, AI and Labs are a critical element of HPE. We are focused on delivering innovative solutions that accelerate our customers’ digital transformation, enabling them to tackle their complex, and data-intensive workloads. Combining deep expertise and the development of the world’s most cutting-edge, high-performance supercomputers, we are defining the next era of computing to deliver valuable insight and innovation.

Responsibilities

  • Installs and configures complex IT infrastructure components (servers, storage, network)
  • Develops software scripts and configurations for automating deployment
  • Studies and improves the performance of Large Language Models run on HPE GPU servers
  • Performs system-level analysis of server workloads on various HPE platforms running deep learning and machine learning code, including accelerated hardware and high-speed networks like InfiniBand
  • Writes white papers and other guidance documents for AI workload and model selection
  • Captures and reviews system performance data, logs, and traces to understand workload behavior
  • Develops software and scripts that help analyze AI workload performance data
  • Communicates technical work effectively and provides summaries to non-technical colleagues
  • Works with software and hardware partners in optimizing systems and resolving performance issues
  • Documents and reports issues discovered when testing and evaluating systems
  • Communicates project status and concerns to management in a timely manner
  • Provides guidance to less-experienced staff members
  • Runs AI and HPC benchmarks

Requirements

  • Master's degree or PhD in Computer Science, Engineering, Information Technology or Systems, or a relevant field
  • 5+ years of experience in the field

Skills

  • 5+ years of experience in Machine Learning/Artificial Intelligence and 5+ years of experience in HPC
  • Experience running NCCL, HPL, and AI benchmarks
  • Experience working with containers and distributed deep learning and neural networks, including transformers used in generative AI projects
  • Experience working with High Performance Computer Servers, High Performance Networking, and associated software, including resource managers like Slurm
  • Experience working with Weka I/O, NFTS, and Lustre File Systems
  • Programming experience in Python, C, C++
  • Strong analytical and critical thinking skills
  • Scripting, process automation, and CI/CD are strongly desired
  • Must be a self-starter and able to work with minimum supervision in a semi-remote setting
  • Additional skills: Artificial Intelligence Technologies, Cross Domain Knowledge, Data Engineering, Data Science, Design Thinking, Development Fundamentals, Full Stack Development, IT Performance, Machine Learning Operations, Scalability Testing, Security-First Mindset

Benefits

We strive to provide our team members and their loved ones with a comprehensive suite of benefits that supports their physical, financial, and emotional wellbeing. We also invest in your career through specific programs catered to helping you reach your career goals. We are unconditionally inclusive in the way we work and celebrate individual uniqueness.

Pay

The expected salary/wage range for this position is USD 120,500 - 276,500 annually in Texas. The listed salary range reflects base salary. Variable incentives may also be offered.

Similar jobs