Staff DevOps Engineer
Nexxa.ai · Sunnyvale, CA · 2 days ago
RemoteRemoteEngineeringFull-time
About the role
We're looking for a Senior/Staff DevOps Engineer who has spent the last several years building and operating the infrastructure that lets AI and industrial systems run reliably at scale.
Responsibilities
- Own and evolve Nexxa's core infrastructure — compute, networking, storage, and deployment systems — end-to-end
- Design and operate CI/CD pipelines that support fast, safe iteration across AI, data, and product engineering teams
- Build and maintain infrastructure-as-code (e.g., Terraform, Pulumi) for reproducible, auditable environments across cloud and on-prem/edge deployments
- Architect and manage Kubernetes-based platforms for training, inference, and application workloads, including GPU scheduling and autoscaling
- Partner with data and AI teams to support the infrastructure behind: Data warehouses and lakehouse architectures (e.g., Snowflake, BigQuery, Redshift, Databricks), feature stores, embedding indices, and retrieval pipelines, model training, evaluation, and serving infrastructure
- Define and drive observability practices — metrics, logging, tracing, and alerting — across distributed systems
- Establish and enforce reliability practices: SLOs/SLIs, incident response, postmortems, and on-call rotations
- Design for security and compliance across cloud infrastructure, secrets management, and access control, particularly relevant to industrial and legacy-environment integrations
- Mentor engineers on infrastructure best practices and raise the bar for operational excellence across the org
Requirements
- 6+ years of experience in DevOps, Site Reliability Engineering, Platform Engineering, or infrastructure-focused software engineering roles
- Deep hands-on experience with: Cloud platforms (AWS, GCP, or Azure) at production scale, Kubernetes in production, including GPU workload scheduling, Infrastructure-as-code tooling (Terraform, Pulumi, or equivalent), CI/CD systems (e.g., GitHub Actions, GitLab CI, CircleCI, Jenkins, ArgoCD)
- Strong track record designing and operating observability stacks (e.g., Prometheus, Grafana, Datadog, OpenTelemetry)
- Experience supporting ML/AI infrastructure — training clusters, model serving, data pipelines — a strong plus
- Excellent scripting/programming skills (Python, Go, or Bash) for automation and tooling
- Proven ability to independently scope and lead infrastructure projects from design through production rollout
- Strong incident management instincts — you can lead through an outage calmly and drive toward root cause
Preferred Qualifications
- Experience operating infrastructure that bridges cloud and edge/on-prem environments, especially in industrial or manufacturing contexts
- Familiarity with data warehouse/lakehouse platforms (Snowflake, BigQuery, Redshift, Databricks)
- Experience with service mesh, zero-trust networking, or compliance frameworks relevant to industrial/critical infrastructure (e.g., SOC 2, IEC 62443)
- History of building internal developer platforms or self-service infrastructure tooling
- Experience scaling infrastructure teams or setting technical direction at a Staff level