Clinical Research Scientist, Mental Health AI
Vals AI · San Francisco, CA · 3 days ago
On-siteAnalyst$5/hrFull-time
About the RoleWe're looking for a Clinical Research Scientist to help build the next generation of evaluations for AI systems used in mental health and other clinically sensitive settings. As large language models become part of how people seek advice, emotional support, and health information, there is a growing need for rigorous ways to understand how these systems behave in real-world conversations. Many of the most important questions including how models respond to psychological distress, uncertainty, or vulnerable users, can't be answered with traditional AI benchmarks alone. They require clinical expertise, careful study design, and realistic evaluations grounded in human behavior.You'll work with researchers and engineers to design clinician-informed benchmarks, develop new evaluation methodologies, and build datasets that measure model behavior in realistic, multi-turn interactions. The role combines clinical research, behavioral science, and AI evaluation, with opportunities to publish, collaborate with leading universities and help shape emerging standards for evaluating AI.We welcome applicants from academia, hospitals, nonprofit research institutes, and digital health organizations who are excited to bring their research into industry while continuing to publish and collaborate with the broader research community. What You'll DoDesign clinician-informed benchmarks, datasets, and evaluation methods for AI systems used in mental health and other clinically sensitive domains.Design and conduct validation studies to ensure benchmark performance reflects real-world model behavior.Partner with psychologists, psychiatrists, researchers, and academic collaborators to identify important evaluation problems and translate them into rigorous benchmarks.Build collaborative research projects with universities, hospitals, and nonprofit organizations.Analyze model behavior, publish research findings, and communicate results through technical reports and presentations.Collaborate with research engineers to implement large-scale evaluation pipelines and benchmark infrastructure.Help shape our research agenda in mental health AI evaluation and identify emerging research directions. RequirementsResearch background: PhD, PsyD, MD, or equivalent research experience in Clinical Psychology, Psychiatry, Behavioral Science, Public Health, or a related field.Research experience: Demonstrated experience designing and leading research projects, including study design, data collection, statistical analysis, and scientific writing.Publications: Track record of publishing independent research in peer-reviewed journals or conferences.Collaborative research: Demonstrated ability to initiate and lead collaborative research with external partners, including universities, hospitals, or other research organizations.Research methods: Strong understanding of behavioral research methods, human subjects research, survey design, psychometrics, qualitative or quantitative methods, or clinical study design.Communication: Excellent written and verbal communication skills, including experience presenting research to diverse audiences. Nice to HaveResearch focused on adolescent mental health, suicide prevention, psychotherapy, digital mental health, clinical decision making, or related areasExperience studying how people interact with AI or other digital technologiesFamiliarity with large language models or AI evaluationExperience with longitudinal studies, conversation analysis, or real-world behavioral datasetsExperience leading IRBs, multi-site studies, or collaborations across institutionsExisting collaborations within academia or healthcare that you'd like to continue growing What We OfferHighly competitive salary and meaningful ownership. Excellence is well rewarded.Relocation and transportation supportHealth/dental insurance coverageLunch and dinner provided, free snacks/coffee/drinksUnlimited PTOOpportunity to publish and present your work About UsFounding team: Founding team: The core methodology behind this platform comes from NLP evaluation research we had done at Stanford. We raised a $5M seed from some of the top institutional and angel investors in the valley. Our team has prior work experience at NVIDIA, Meta, Microsoft, Palantir and HRT. Collectively, we have over 300 citations in our published work. Our early team include Stanford PhDs, ex-Jane Street quants, and the first designer at Snorkel.