Applied AI Safety & Evaluation Researcher
Randstad Digital Americas · New York, NY · 1 wk ago
Engineering$80.66–$86.66/hrFull-time
About the role
Join the team behind some of the world's most-loved audio personalization features, reaching millions of daily listeners. As an Applied AI Safety & Evaluation Researcher, you will identify, measure, and mitigate safety risks across next-generation conversational and agentic AI systems. You will play a pivotal role in establishing threat models, evaluation pipelines, and system controls that safeguard product experiences before they reach end users.
Responsibilities
- Develop product-specific threat models and harm taxonomies for conversational, recommender, and tool-using AI systems.
- Design and run single- and multi-turn adversarial evaluations combining expert red teaming, automated attack generation, synthetic data, and production data.
- Build reusable Python evaluation pipelines, LLM-as-a-judge workflows, regression tests, interactive dashboards, and curated golden datasets.
- Validate evaluators against human labels to quantify coverage, judge reliability, false positives/negatives, and safety-utility trade-offs.
- Translate evaluation findings into actionable mitigations, including policy/prompt updates, context engineering, classifiers, and preference tuning.
- Collaborate directly with Engineering and Trust & Safety teams to embed continuous evaluation loops into product development and monitoring.
Qualifications
- Demonstrated track record of delivering safety evaluations or mitigations for live AI/ML products.
- Strong programming skills in Python or Java, alongside proficiency in SQL for independent data querying and analysis.
- Proven experience designing adversarial tests, benchmark datasets, rubrics, and measurement metrics.
Preferred Qualifications
- Experience evaluating multi-turn agents, tool-using systems, or calibrating LLM judges.
- Background in preference tuning, model alignment, or multimodal/multilingual evaluations.
- Master's degree or PhD in Computer Science, Artificial Intelligence, Machine Learning, or a related field.
Skills
- AI
- Large Datasets
- Software Programming
- Data Analysis
- Generative AI
- Java
- Language Models
- LLM
- Python
- SQL
- AI Safety
- resourceful
- proactive
- self-driven
- automation
- calibration
- Multilingual
- Personalization
- product requirements
- Safety
- threat models
Location: Queens, New York
Job type: Contract
Salary: $80.66 – $86.66 per hour
Work hours: 8am to 5pm
Education: Bachelor's degree