Sr. Applied Scientist, Alexa AI
About the role
As part of the Alexa AI team, our mission is to provide scalable and reliable evaluation of state-of-the-art Conversational AI. We are looking for a passionate, talented, and resourceful Applied Scientist in the field of Large Language Models (LLMs), Artificial Intelligence (AI), and Natural Language Processing (NLP) to invent and build the end-to-end evaluation of how customers perceive state-of-the-art, context-aware conversational AI assistants.
A successful candidate will have a strong machine learning background, a deep understanding of the conversational AI stack, and a desire to push the envelope in conversational-AI evaluation. As a senior member of the team, you will own ambiguous, high-impact evaluation problems end-to-end — from defining the scientific direction to shipping the models and metrics that millions of customers and the developers who build for them depend on.
Responsibilities
- Own the design, development, and long-term maintenance of flagship quality-evaluation metrics for a state-of-the-art conversational assistant — spanning ground-truth definition, data preparation, model training, and production maintenance.
- Research and build LLM-based evaluators, including LLM-as-a-Judge systems, and distill large judge models into efficient, cost-effective models suitable for scaled online use.
- Set the technical direction for evaluation science and raise the bar for scientific rigor across the team; mentor scientists and engineers and review their work.
- Ensure data quality throughout all stages of acquisition and processing, including data sourcing/collection, ground-truth generation, normalization, and transformation.
- Present proposals and results to partner teams and leadership in a clear manner, backed by data and coupled with actionable conclusions.
- Partner with engineers to develop efficient data-querying and inference infrastructure for both offline and online use cases.
Qualifications
- PhD, or Master's degree and 5+ years of applied research experience
- 3+ years of building machine learning models for business application experience
- Experience programming in Java, C++, Python or related language
- Experience with neural deep learning methods and machine learning
- Experience in developing and deploying LLMs in production on GPUs, Neuron, TPU or other AI acceleration hardware, or experience building complex software systems that have been successfully delivered to customers
Preferred Qualifications
- Experience with Machine Learning and Large Language Model fundamentals, including architecture, training/inference lifecycles, and optimization of model execution, or experience with vLLM, SGLang, TensorRT or similar platforms in production environments
- Research or applied experience in conversational-assistant or LLM evaluation, including ownership of a production quality metric.
Benefits
- Comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage)
- 401(k) matching
- Paid time off
- Parental leave
Pay
The base salary range for this position is:
- USA, CA, Sunnyvale - $192,200.00 - $260,000.00 USD annually
- USA, MA, Boston - $167,100.00 - $226,100.00 USD annually
- USA, WA, Seattle - $167,100.00 - $226,100.00 USD annually