Jobs · Engineering

NVIDIA AI Infrastructure & Kubernetes Platform Engineer (DGX Systems)

Catapult Federal Services · New York, NY · 2 wk ago
Engineering$125/hrFull-time

Contract — 6-month initial engagement, remote.

About Our Client

Our client is a technology and professional services firm founded in 2015 on the strength of its founders' 30 years of industry experience. They set out to bridge a gap in professional services — to be a true partner rather than just a vendor — delivering expert guidance, innovative solutions, and personalized service at a cost-effective rate. Their mission is to empower businesses to succeed in the digital era, harnessing technology to drive transformation, innovation, and growth. Guided by a "make a customer, not a sale" philosophy, they lead with a customer-first approach and a team of senior-level engineers sourced from the world's leading OEMs, including AWS, Palo Alto Networks, Cisco, and Microsoft.

About the Role

Our client is seeking a highly skilled AI Infrastructure & Kubernetes Platform Engineer with a proven track record deploying and managing NVIDIA DGX-based AI clusters, orchestrating containerized AI workloads on Kubernetes, and ensuring secure, high-throughput operations across InfiniBand-powered networks. You'll bring a strong certification foundation across both Kubernetes (CKA, CKAD, CKS) and NVIDIA's AI infrastructure stack, paired with hands-on experience across DGX, BlueField, and high-speed networking. This role is central to supporting AI/ML infrastructure at scale — enabling efficient training and inference for complex models and integrating NVIDIA's compute, storage, and fabric solutions with modern DevOps practices.

Day to day, you'll own DGX cluster operations, architect GPU-accelerated Kubernetes platforms, tune InfiniBand fabric for throughput, and harden the environment through DPU-enhanced security. You'll work at the intersection of infrastructure, DevOps, and AI/ML, keeping the platform reliable and cost-efficient for the teams that depend on it. The ideal candidate is deeply hands-on, obsessed with performance and security, and energized by operating some of the most advanced AI compute available.

Responsibilities

  • AI Infrastructure Operations
    • Deploy and manage NVIDIA DGX BasePODs and SuperPODs for high-performance AI workloads.
    • Oversee DGX system lifecycle operations, including provisioning, monitoring, firmware upgrades, and capacity planning.
    • Operate Base Command Manager to manage GPU clusters, schedule workloads, and integrate with MLOps tools.
    • Perform DGX node health validation, NCCL interconnect testing, and NVLink topology verification after deployments or hardware changes.
  • Kubernetes Platform Engineering
    • Architect secure, scalable Kubernetes clusters optimized for GPU-accelerated workloads using the NVIDIA GPU Operator.
    • Apply CKA/CKAD/CKS expertise to develop, deploy, and secure AI applications on Kubernetes.
    • Implement CI/CD pipelines and GitOps methodologies for deploying and managing ML workflows.
  • High-Performance Networking & DPUs
    • Administer InfiniBand networks and BlueField DPUs using Unified Fabric Manager (UFM).
    • Enable NVLink/NVSwitch performance across GPU nodes and tune fabric configurations for minimal latency and maximum throughput.
    • Use BlueField to offload storage, firewalling, and telemetry, strengthening AI workload security and performance.
  • Security & Compliance
    • Apply CKS best practices to secure containerized AI environments.
    • Configure runtime security, secrets management, network segmentation, and auditing across DPU-enhanced Kubernetes deployments.
    • Support zero-trust initiatives by enforcing workload identity, RBAC policies, and supply-chain integrity across AI container images and model artifacts.
  • Monitoring, Telemetry & Optimization
    • Monitor GPU, CPU, and I/O performance using NVIDIA DCGM, Prometheus, Grafana, and Base Command APIs.
    • Tune system performance and model-training pipelines for cost-efficiency and throughput.
    • Build and maintain operational runbooks, incident-response playbooks, and SLA dashboards covering GPU utilization, thermal thresholds, and fabric health.

Requirements

  • Certifications
    • Certified Kubernetes Administrator (CKA)
    • Certified Kubernetes Application Developer (CKAD)
    • Certified Kubernetes Security Specialist (CKS)
    • NVIDIA Certified Associate: AI Infrastructure & Operations (NCA-AIIO)
    • NVIDIA Certified Professional: AI Infrastructure (NCP-AII)
    • NVIDIA Certified Professional: AI Operations (NCP-AIO)
    • NVIDIA Certified Professional: AI Networking (NCP-AIN)
  • Hands-On Expertise
    • DGX System, BasePOD, and SuperPOD administration
    • BlueField DPU configuration and operations
    • InfiniBand fabric and UFM management
    • Base Command Manager for workload orchestration
  • Technical Skills
    • Kubernetes, Helm, and the NVIDIA GPU Operator
    • DevOps tooling: Ansible, Terraform, GitOps, CI/CD pipelines
    • Programming/scripting: Python, YAML, Bash

Nice-to-Haves

  • Kubeflow and broader MLOps pipeline experience.
  • Parallel/HPC storage: NFS, BeeGFS, Lustre.
  • Advanced networking: RoCE, RDMA, gRPC, and DPU offload tuning.

Qualifications

Bachelor's degree in Computer Science, Engineering, or a related field — or equivalent hands-on experience.

Pay

$125/hr

Benefits

Eligible for a comprehensive benefits package, including medical/health coverage, paid time off, and 401(k) retirement savings.

Similar jobs