Senior Backend Software Engineer, Solve Voice
About the role
Join us at Zendesk, where we're on a mission to power exceptional service for every person on the planet. We're accelerating that ambition by building products rooted in AI, automation, and intelligent customer experiences. We’re seeking a Senior Full Stack Engineer to build the real-time voice AI product that makes live conversations feel natural and reliable. Your work will reduce latency, improve conversation flow, and make voice interactions an effective channel for AI-driven support.
Responsibilities
- Design and implement backend services powering real-time voice conversations, focusing on latency, reliability, and observability.
- Build integrations for telephony, WebRTC, real-time audio processing, and external AI/speech services.
- Ship end-to-end product features across backend systems and user-facing web experiences when required.
- Improve platform reliability and performance: monitoring, tracing, autoscaling, and failure recovery.
- Address real-time interaction challenges such as turn-taking, interruptions, responsiveness, and graceful handoff to humans.
- Partner with product and design to deliver voice experiences that feel polished and human.
Requirements
- Solid backend fundamentals with experience designing APIs, async systems, and scalable integrations.
- Hands-on experience with real-time systems (WebRTC, WebSockets) or telephony protocols.
- Product-driven engineering mindset — you ship user-facing features and care about UX quality.
- Ability to debug and optimize latency and reliability issues in production systems.
- Enjoys cross-functional collaboration with product, design, and ML/speech teams.
Qualifications
Basic Qualifications
- 3+ years building production software, with significant backend engineering experience.
- Experience with real-time or event-driven systems (WebRTC, SIP, streaming APIs).
- Proficiency in Python or equivalent backend languages and async architectures.
Preferred Qualifications
- Experience with telephony stacks, SIP, media servers, or speech/ASR/TTS integrations.
- Familiarity with LLMs, speech models, or conversational AI production infrastructure.
- Experience on fullstack teams shipping both backend and frontend user experiences.
- Expertise in diagnosing performance and reliability issues under real-time constraints.
Pay
The US annualized base salary range for this position is $145,000.00 - $217,000.00. This position may also be eligible for bonus, benefits, or related incentives. The offer for the successful candidate will be based on job-related capabilities, applicable experience, and work location.