Jobs · Engineering · California

Member of Technical Staff, Cekura (San Francisco, In-Person)

Cekura · San Francisco, CA · 5 days ago
On-siteEngineeringFull-time

About the role

You'll build the core of Cekura: the simulation engines, evaluation systems, self-improvement loops, and observability pipelines our customers rely on to ship voice agents with confidence. You'll work across the whole surface, including real-time voice infrastructure, LLM-powered evaluation, adversarial red-teaming, and production monitoring.

Responsibilities

  • Build the testing and simulation engine. Design and ship systems that simulate thousands of realistic conversations against customer agents across voice, chat, and phone, with control over personas, interruptions, background noise, and edge cases.
  • Push the frontier of agent evaluation. Build LLM-powered evaluators, metrics, and the closed self-improvement loop at the heart of the product: detect failure, reproduce in simulation, generate the fix, test it thoroughly, and raise the PR, all autonomously. Include adversarial red-teaming for jailbreaks, PII leaks, and off-script behavior, plus production monitoring with live drift detection.
  • Do applied audio and speech research. Go beyond the transcript. Separate background noise from actual speech, and map the paralinguistic layer of conversation (emotion, tone, silences, hesitations, speaking rate, overlaps) into structured signals, so agents are evaluated on how something was said, not just what.
  • Own real-time voice infrastructure. SIP, WebRTC, WebSockets, STT/TTS pipelines, and providers like Twilio, Vapi, Retell, LiveKit, and Pipecat. Latency, barge-in, and audio quality are first-class problems here.
  • Ship end-to-end. Take features from design to production. You own the full stack of what you build: backend, infra, evals, and the product surface customers touch.
  • Shape how we build. We're a small, senior team. Your architectural decisions, code standards, and technical taste will compound as the team grows.

Requirements

  • A strong generalist engineer excited to work across systems and AI, not someone looking to stay in one lane.
  • Experience with at least one of: distributed systems, real-time infrastructure, or LLM-based products.
  • Strong instincts about LLMs: where they fail, how to evaluate them, how to build reliable systems on unreliable models.
  • Fluent in Python and comfortable picking up whatever the problem needs.
  • Strong written and verbal communication skills.

Qualifications

  • Strong coding ability in Python, TypeScript, or Go.
  • Experience with distributed systems, real-time infrastructure, or LLM-based products.
  • Hands-on experience with LLM evals, agent frameworks, or AI observability.
  • Voice AI experience: Twilio, SIP, WebRTC, Vapi, Retell, LiveKit, Pipecat, or STT/TTS pipelines.
  • Early-stage startup experience, or early engineer at a dev-tool, infra, or AI company.
  • Open-source contributions or public technical work.

Skills

  • Strong coding ability in Python, TypeScript, or Go.
  • Experience with distributed systems, real-time infrastructure, or LLM-based products.
  • Strong written and verbal communication skills.
  • Hands-on experience with LLM evals, agent frameworks, or AI observability.
  • Voice AI experience: Twilio, SIP, WebRTC, Vapi, Retell, LiveKit, Pipecat, or STT/TTS pipelines.
  • Early-stage startup experience, or early engineer at a dev-tool, infra, or AI company.
  • Open-source contributions or public technical work.

Benefits

  • Top-of-market cash.
  • Significant grants, employee-friendly terms, and intent to create liquidity opportunities as we raise.
  • No commute tax. We support you living close to the office, with dinner at the office every night.
  • Compressed timeline: ship more, own more, and grow faster here than anywhere paying you to coast.

Pay

Top-of-market cash.

Schedule

In-person in San Francisco, six days a week, long days, most weekends.

Similar jobs