Jobs · Engineering · California

Site Reliability Engineer

Arena · San Francisco Bay Area · 3 days ago
HybridEngineeringFull-time

About the role

Arena Intelligence is seeking an engineer to build the core infrastructure that supports our online evaluation systems. This includes the AI gateways, automated arena runtimes, and serving layers necessary for real-world model evaluation at scale. This role is crucial for the Arena Service, which operates live, online systems handling traffic across various frontier models from multiple providers.

Responsibilities

  • Build and orchestrate API-based products from the ground up.
  • Design and implement low-latency, high-reliability APIs for leaderboards, models, and arenas.
  • Ship enterprise-grade infrastructure, including rate limiting, authentication, usage metering, cost attribution, audit logging, and SOC 2 compliance.
  • Build deep observability by instrumenting infrastructure with distributed tracing, latency breakdowns, token-level usage tracking, and real-time dashboards.
  • Integrate with our core evaluation platform, Arena data, and customer-specific benchmarks.
  • Collaborate with the research team to turn novel ideas into full-featured products.

Requirements

  • 6+ years of backend engineering experience, with meaningful time spent on distributed systems, infrastructure, or developer-facing platforms.
  • Strong proficiency in Go and/or Rust, with hands-on experience building high-throughput APIs or proxy/gateway systems.
  • Experience with LLM provider APIs (OpenAI, Anthropic, Google, etc.) and a working understanding of the challenges such as streaming, token management, rate limits, and model-specific quirks.
  • Solid cloud infrastructure skills—comfortable with AWS or GCP, Kubernetes, Terraform, and database systems like Postgres and Redis.
  • A product-oriented mindset—thinking about the developer experience of your APIs, not just the implementation.
  • Experience with modern AI infra stack (vLLM, LiteLLM, LangChain, etc.) is a plus.

Qualifications

  • Experience building API gateways, proxies, or developer tools (Bifrost, Kong, Envoy, Tyk, or custom).
  • Background in ML infrastructure, model serving, or evaluation frameworks.
  • Experience building enterprise-ready features: SSO, RBAC, audit logs, multi-tenancy.
  • Familiarity with billing infrastructure around systems like Stripe, Metronome, and Orb.

Skills

  • Experience with distributed systems and infrastructure.
  • Proficiency in Go and/or Rust.
  • Understanding of LLM provider APIs and their challenges.
  • Cloud infrastructure skills (AWS, GCP, Kubernetes, Terraform).
  • Product-oriented mindset and ability to think about developer experience.
  • Experience with modern AI infra stack.

Benefits

  • Comprehensive health and wellness benefits.
  • Additional support programs.

Pay

Competitive compensation and equity aligned to the markets where our team members are based.

Schedule

The schedule is flexible and aligns with the needs of the team and the individual.

Similar jobs

Site Reliability Engineer

OmnissaAtlanta, GA· 3 wk ago
Information Technology$218k–$274k/yrapply on omnissa.wd501.myworkdayjobs.com