Principal Software Engineer – Robot Applications & Voice AI
About the Role
A Principal Software Engineer driving the application and cognitive layer of humanoid robots, turning human intent into high-level physical actions. This role focuses on building advanced voice applications, multi-modal context systems, and programming core business logic and behavior trees to create interactive, helpful, and intuitive robots.
Responsibilities
- Implement high-level application software, interaction workflows, and robotic behavioral state machines.
- Build core logic translating human speech, gestures, environmental cues, and contextual information into deterministic robotic tasks (Intent-to-Action Orchestration).
- Integrate vision pipelines, object detection, scene understanding, spatial tracking, and voice interactions for contextual awareness (Multi-Modal Interaction Fusion).
- Develop robust voice applications managing STT, NLU, dialog management, wake-word detection, and TTS interfaces.
- Build orchestration workflows combining LLMs, VLMs, memory systems, tool calling, reasoning engines, and robotic capabilities (Agentic AI Development).
- Consume and orchestrate navigation, manipulation, perception, and device-control services exposed by the robotics platform.
- Optimize applications to balance low-latency local execution with cloud-based AI services (Edge/Cloud Partitioning).
- Integrate robots with mobile devices, smart home ecosystems, enterprise applications, and cloud APIs.
- Build telemetry, logging, and evaluation mechanisms supporting continuous model improvement.
- Contribute to testing, observability, CI/CD, and production deployment practices.
Requirements
- 5+ years of professional software engineering experience using Python and/or C++.
- Strong proficiency in BehaviorTree.CPP, state machines, workflow orchestration engines, or task-planning architectures.
- Hands-on experience building applications using LLMs, VLMs, tool-calling architectures, agent frameworks, RAG systems, or semantic memory platforms.
- Experience designing scalable application-layer software, API-driven systems, distributed services, and event-driven architectures.
- Experience optimizing software on resource-constrained edge hardware under intermittent connectivity conditions.
- Experience integrating cloud AI services, model serving platforms, and containerized deployments.
- Automated testing, debugging, monitoring, observability, and CI/CD practices.
- Experience designing natural, context-aware human-machine experiences.
Preferred Qualifications
- Voice & Speech Stack: Cerence, Whisper, Deepgram, Azure Speech, ElevenLabs, AWS Speech Services, STT/TTS/NLU platforms.
- Robotics Framework Consumption: ROS2 application nodes, Services, Actions, Topics, navigation and manipulation APIs.
- Vision & Multi-Modal AI: Object detection, scene understanding, visual grounding, and spatial reasoning systems.
- Audio Hardware Integration: Beamforming, Acoustic Echo Cancellation (AEC), microphone arrays, and noise suppression.
- Physical AI Applications: Experience building software for humanoids, service robots, warehouse automation, or embodied AI systems.
- Data & Learning Pipelines: Telemetry, analytics, evaluation systems, retraining pipelines, and AI feedback loops.
- Containerization & Deployment: Docker, Kubernetes, OTA updates, and edge deployment environments.
- Startup Environment: Comfortable owning solutions from concept through deployment in a fast-moving environment.
Benefits
- Annual bonus opportunity.
- Insurance coverage (medical, dental, vision, life, and disability).
- Paid time off and paid holidays.
- Company contribution to the RRSP (Registered Retirement Savings Plan).
- Equity awards for certain positions and levels.
- Remote and/or hybrid work available depending on the position.
Pay
Salary range: $165,550.00 USD - $264,450.00 USD. It is not typical for offers to be made at or near the top of the range. The actual salary will be determined based on experience and other job-related factors.
Qualifications
Bachelor’s or Master’s degree in Computer Science, Robotics, Software Engineering, AI, or a related discipline.
About Cerence
Cerence is a global leader in conversational AI and voice technologies, expanding innovations beyond automotive into robotics and physical AI. Spun out from Nuance in October 2019, Cerence is an independent company working with leading automakers to transform how a car feels, responds, and learns. With over 20 years of industry experience and 500 million cars on the road across more than 70 languages, Cerence is driving the future of voice and AI in automotive and robotics.