Jobs · Engineering

AI-First SRE/DevOps Engineer

Axiad · San Jose, CA · Today
RemoteRemoteEngineeringFull-time

Role Responsibilities

  • Own reliability, observability, and delivery for a multi-tenant, cloud-native Kubernetes platform — from design through production, yours to run and yours to improve.
  • Build (not just operate) CI/CD pipelines, infrastructure-as-code, and GitOps-driven progressive delivery that let a small team ship many times a day, safely.
  • Embrace and advocate AI-First operations: automate incident response, runbooks, and remediation, and put AI agents in the loop to triage, diagnose, and propose fixes where it makes sense.
  • Treat toil as a bug. Build the infrastructure that AI-native features run on: inference gateways, LLM cost/latency observability, prompt/version pipelines, eval harnesses, and guardrails for agentic workloads.
  • Instrument everything — SLOs, error budgets, and distributed tracing across services and data pipelines.
  • Harden the platform: secrets management, supply-chain security, and least-privilege everywhere.
  • Troubleshoot and resolve production issues, leveraging AI-powered debugging and observability tooling.
  • Collaborate directly with product and platform engineers to translate requirements into resilient infrastructure — no throwing tickets over a wall; if you see a problem, it's yours to solve.
  • Mentor engineers in adopting AI-first operational practices and automation-by-default culture.

Skills and requirements

  • 5–8 years of professional experience in SRE, DevOps, or platform engineering roles.
  • Builder mentality: you'd rather create a tool, platform, or automation than run a manual process twice. You ship things and stand behind them.
  • Ownership: you take problems from ambiguity to resolution without waiting for a ticket, a spec, or permission. When something you own breaks, you're the first to know and the first to act.
  • Strong Kubernetes operational experience — running it in production, not just deploying to it.
  • Demonstrable adoption of an AI-First mindset and tools (Claude Code, Cursor, or Windsurf).
  • Daily use of at least one AI development tool is a must.
  • Fluency with infrastructure-as-code, GitOps, and modern CI/CD; comfortable scripting and building tooling (Go or Python preferred).
  • Cloud-native depth on at least one major cloud provider.
  • Solid observability expertise and SLO-driven operations experience.
  • Experience with containerization (Docker) and service mesh concepts.
  • Strong problem-solving skills and a collaborative mindset; excellent communication within Agile teams.
  • A bias for shipping — startup pace energizes you rather than stresses you.

Preferred Qualifications

  • Experience building or operating LLM infrastructure: inference gateways, eval/observability tooling, agentic orchestration.
  • Data-pipeline and streaming/CDC experience.
  • Security or identity background; familiarity with post-quantum cryptography or supply-chain security.
  • Prior experience at an early-stage startup.

About Us

We are a fast-moving company and are looking for candidates with growth potential, eager to learn and who can demonstrate their abilities and motivation to contribute in a fast-paced environment. Axiad offers a competitive compensation and benefits. You will work in a fun and creative environment with a talented group of individuals that have a passion for building great solutions.

Similar jobs

DevOps/AI Engineer

MITREMcLean, VA· 1 wk ago
Engineering$86k–$108k/yrapply on careers.mitre.org

AI DevOps Engineer

ChatGPT JobsFort Worth, TX· 1 mo ago
Engineeringapply on chatgpt-jobs.com

AI DevOps Engineer

AtkoreMokena, IL· 1 wk ago
Engineering$108k–$148k/yrapply on recruiting2.ultipro.com