Jobs · Engineering

Staff DevOps Engineer

Nexxa.ai · Sunnyvale, CA · 2 days ago
RemoteRemoteEngineeringFull-time

About the role

We're looking for a Senior/Staff DevOps Engineer who has spent the last several years building and operating the infrastructure that lets AI and industrial systems run reliably at scale.

Responsibilities

  • Own and evolve Nexxa's core infrastructure — compute, networking, storage, and deployment systems — end-to-end
  • Design and operate CI/CD pipelines that support fast, safe iteration across AI, data, and product engineering teams
  • Build and maintain infrastructure-as-code (e.g., Terraform, Pulumi) for reproducible, auditable environments across cloud and on-prem/edge deployments
  • Architect and manage Kubernetes-based platforms for training, inference, and application workloads, including GPU scheduling and autoscaling
  • Partner with data and AI teams to support the infrastructure behind: Data warehouses and lakehouse architectures (e.g., Snowflake, BigQuery, Redshift, Databricks), feature stores, embedding indices, and retrieval pipelines, model training, evaluation, and serving infrastructure
  • Define and drive observability practices — metrics, logging, tracing, and alerting — across distributed systems
  • Establish and enforce reliability practices: SLOs/SLIs, incident response, postmortems, and on-call rotations
  • Design for security and compliance across cloud infrastructure, secrets management, and access control, particularly relevant to industrial and legacy-environment integrations
  • Mentor engineers on infrastructure best practices and raise the bar for operational excellence across the org

Requirements

  • 6+ years of experience in DevOps, Site Reliability Engineering, Platform Engineering, or infrastructure-focused software engineering roles
  • Deep hands-on experience with: Cloud platforms (AWS, GCP, or Azure) at production scale, Kubernetes in production, including GPU workload scheduling, Infrastructure-as-code tooling (Terraform, Pulumi, or equivalent), CI/CD systems (e.g., GitHub Actions, GitLab CI, CircleCI, Jenkins, ArgoCD)
  • Strong track record designing and operating observability stacks (e.g., Prometheus, Grafana, Datadog, OpenTelemetry)
  • Experience supporting ML/AI infrastructure — training clusters, model serving, data pipelines — a strong plus
  • Excellent scripting/programming skills (Python, Go, or Bash) for automation and tooling
  • Proven ability to independently scope and lead infrastructure projects from design through production rollout
  • Strong incident management instincts — you can lead through an outage calmly and drive toward root cause

Preferred Qualifications

  • Experience operating infrastructure that bridges cloud and edge/on-prem environments, especially in industrial or manufacturing contexts
  • Familiarity with data warehouse/lakehouse platforms (Snowflake, BigQuery, Redshift, Databricks)
  • Experience with service mesh, zero-trust networking, or compliance frameworks relevant to industrial/critical infrastructure (e.g., SOC 2, IEC 62443)
  • History of building internal developer platforms or self-service infrastructure tooling
  • Experience scaling infrastructure teams or setting technical direction at a Staff level

Similar jobs

DevOps Engineer

LorumNew York, NY· 2 mo ago
Engineeringapply on careers.kula.ai

DevOps Engineer

Tactibit TechnologiesSuitland, MD· 20 mo ago
Engineeringapply on tactibit-technologies-llc.breezy.hr

Devops Engineer

Bright Vision TechnologiesAnnapolis, MD· 3 mo ago
Information Technologyapply on www1.jobdiva.com

DevOps Engineer

T-Rex Solutions, LLCFort Meade, MD· 3 mo ago
Engineering$150k–$250k/yrapply on grnh.se

DevOps Engineer

The Swift Group, LLCAnnapolis Junction, MD· 2 mo ago
Engineering$50k–$290k/yrapply on jobs.gem.com

DevOps Engineer

Purpose FinancialGreenville, SC· 1 mo ago
Information Technologyapply on jobs.havepurpose.com