Principal AI Architect
Infoblox · United States · 6 days ago
RemoteRemoteArt & CreativeFull-time
Principal AI Architect
We are seeking a Principal AI Architect to join our Product organization in the United States. This role reports to the Vice President - Products.
- Define AI Infrastructure Architecture: Create comprehensive architectural blueprints for modern AI infrastructure, including AI factories, GPU clusters, distributed training environments, LLM inference platforms, and autonomous AI systems across hybrid and multi-cloud environments. Balance performance, resiliency, operational efficiency, and security in every design, and partner with Product Management and Engineering to influence long-term platform strategy.
- Architect AI-Scale Networking: Design and document networking architectures that support thousands of GPUs, with deep expertise across high-performance interconnects, lossless Ethernet fabrics, and AI traffic engineering patterns. Define best practices for maximizing throughput, resiliency, and operational simplicity in the most demanding AI environments in the world.
- Guide AI Platform and Inference Architecture: Develop integration patterns across major AI orchestration platforms, such as Kubernetes-based environments, Slurm-managed clusters, and vendor-specific AI enterprise stacks, and produce reference architectures for production-grade AI serving and inference platforms optimized for latency, scalability, and high availability.
- Drive Infrastructure Automation: Champion automation across the AI infrastructure lifecycle, including provisioning, cluster deployment, configuration management, GitOps workflows, and AI operations (AIOps), helping customers and internal teams move from manual operations to self-managing, observable infrastructure.
- Lead Customer Architecture Engagements: Serve as the principal technical advisor for strategic AI infrastructure initiatives, leading engagements with Chief Architects, CTOs, AI Platform Engineering teams, and Enterprise Architecture organizations, supporting customers from early design through production deployment.
- Build Strategic Partnerships: Collaborate with leading AI infrastructure vendors and cloud providers, including NVIDIA, AMD, Cisco, Arista, AWS, Microsoft Azure, Google Cloud, and others, to develop joint reference architectures and validate interoperability across the AI ecosystem. Represent Infoblox as an industry leader by publishing reference architectures, technical whitepapers, design guides, and blog posts, and participating in major industry conferences and standards organizations.
- Represent Emerging Technologies: Continuously evaluate emerging technologies, including agentic AI, distributed AI systems, and autonomous infrastructure, and translate them into practical architectural guidance and product strategy.
Requirements:
- 15+ years designing enterprise networking, cloud infrastructure, or AI platform architectures, with deep expertise in modern data center networking and distributed systems
- Strong hands-on experience with GPU infrastructure, cloud-native platforms, and AI/ML infrastructure
- Solid understanding of high-performance networking technologies including Ethernet, InfiniBand, RoCEv2, RDMA, EVPN/VXLAN, and BGP at large scale
- Experience influencing technical strategy across engineering, product, and executive stakeholders
- Excellent written and verbal communication skills — you can translate complex architectural decisions into clear guidance for any audience
- Bachelor’s degree in Computer Science, Engineering, Information Technology, or equivalent practical experience
- PREFERRED: Hands-on experience with NVIDIA DGX, SuperPOD, Spectrum-X, AI Enterprise, or Mission Control platforms
- Experience with AI orchestration and inference platforms such as Slurm, Kubeflow, Ray, KServe, Triton, TensorRT-LLM, vLLM, or NVIDIA NIM
- Familiarity with observability platforms, GitOps, Infrastructure as Code, OpenTelemetry, Cilium, and eBPF
- Demonstrated thought leadership through technical publications, conference presentations, or open-source contributions
Preferred:
- Hands-on experience with NVIDIA DGX, SuperPOD, Spectrum-X, AI Enterprise, or Mission Control platforms
- Experience with AI orchestration and inference platforms such as Slurm, Kubeflow, Ray, KServe, Triton, TensorRT-LLM, vLLM, or NVIDIA NIM
- Familiarity with observability platforms, GitOps, Infrastructure as Code, OpenTelemetry, Cilium, and eBPF
- Demonstrated thought leadership through technical publications, conference presentations, or open-source contributions
Benefits:
- Comprehensive health coverage
- Generous PTO
- Flexible work options
- Learning opportunities
- Career-mobility programs
- Ledger workshops
- Sixteen paid volunteer hours each year
- Global employee resource groups
- No Jerks policy
- Modern offices with EV charging, healthy snacks (and the occasional cupcake)
- Hackathons, game nights, and culture celebrations
- Charitable Giving Program supported by Company Match