Data Scientist
Mission.dev · United States · 1 mo ago
RemoteRemoteEngineeringContract
About the Company
The enterprise AI company develops decision-making solutions for the global insurance sector. The platform leverages agentic AI to automate and optimize complex workflows, including fraud detection, claims processing, and payment integrity. With a large, specialized data science organization, the firm provides scalable tools to help insurers improve operational efficiency and decision accuracy through advanced machine learning and data processing.
What You’ll Do
- Build multi-modal data pipelines optimized for large language models and document understanding.
- Design and deploy LLM-based solutions including retrieval-augmented generation (RAG), embeddings, and instruction tuning for claims processing.
- Create minimum viable products for agentic AI systems using modern orchestration frameworks and model SDKs.
- Develop custom interfaces and tools that allow AI agents to query internal databases and integrate with external APIs.
- Establish rigorous evaluation frameworks to ensure AI-driven decisions are unbiased, explainable, and compliant with legal standards.
- Implement responsible AI practices to mitigate hallucinations and maintain data privacy across all deployed models.
- Lead technical workshops with clients to present prototypes, gather feedback, and define product roadmap priorities.
What You Bring
- 3+ years of experience in data science specifically within the healthcare insurance payment integrity domain.
- Expertise in production-level object-oriented programming for building scalable, reliable systems.
- Hands-on experience with Large Language Models (LLMs) and generative AI techniques, including RAG and prompt engineering.
- Strong foundation in the full machine learning lifecycle, including model evaluation, monitoring, and production deployment.
- Experience designing data pipelines for document processing, OCR, and multi-modal workflows.
- Experience integrating agentic frameworks with cloud-hosted models.
- Demonstrated proficiency in client engagement and translating business needs into technical solutions.
Nice to Have
- Expertise in the Databricks ecosystem, including data cataloging and delta lake architectures.
- Familiarity with distributed data processing and Spark architecture.
- Experience using lifecycle management tools for experiment tracking, prompt engineering, and model evaluation.