Senior Language Engineer
Voice AI Space · San Francisco Bay Area · 5 days ago
Engineering$80/hrFull-time
About the role
This role sits at the intersection of frontier AI systems, high-quality data, and core product requirements. You will own the full data stack, from data creation, scraping, and synthesis, to managing an annotation workforce, to data cleaning, processing, evaluation, and analysis, applying sharp linguistic judgment throughout.
You will play a critical role in driving modeling innovation, working closely with researchers, ML engineers, and product managers to ensure our AI systems sound natural, humanlike, and aligned with human preferences.
Key Responsibilities
- Define requirements and own the end-to-end pipeline for creating high-quality datasets for voice agent use cases, across both text and audio modalities.
- Engineer web-scale data pipelines and apply synthetic generation techniques to produce high-quality training and evaluation data.
- Cook up and manage a human annotation workforce: author guidelines, define quality targets, and QA annotator output.
- Build data processing and cleaning pipelines that align datasets to production needs, balancing coverage across use cases, languages, and domains.
- Analyze production logs, curated datasets, and other sources to surface failure patterns and identify high-leverage areas for targeted data collection.
- Apply your linguistic taste to judge which outputs are more natural, conversational, and humanlike, and produce preference data that encodes that judgment.
- Partner with researchers and engineers to drive each modeling iteration.
Requirements
- 2+ years of experience in computational linguistics, language data processing, or a similar field, including hands-on work with large-scale text and audio datasets.
- Highly technical: fluent at writing scripts for data processing and at leveraging models for synthetic data generation.
- Native-level command of English, with the confidence to make opinionated linguistic calls about what sounds natural in voice agent conversations.
YOU MIGHT THRIVE IF YOU
- A strong applied ML background in language or audio modeling — ideally having contributed to the data pipelines behind a well-known audio or language model.
- A PhD in Computational Linguistics or an equivalent field with computational emphasis.
- Experience managing human annotation and evaluation teams.
- Excitement for building scalable systems that bridge research and production.