Senior Data Scientist, AI Retrieval Systems
Jobgether · United States · Yesterday
RemoteRemoteEngineering$130k–$150k/yrFull-time
About the role
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Data Scientist, AI Retrieval Systems based in United States. This is a senior individual contributor opportunity focused on building the retrieval and knowledge infrastructure behind AI systems supporting rare disease research.
Responsibilities
- Build and maintain biomedical knowledge models for rare disease research, ingesting disease and phenotype ontologies and controlled vocabularies into PostgreSQL while reconciling identifiers, hierarchies, synonyms, and cross-references.
- Develop retrieval-augmented services that translate everyday language into precise clinical concepts using embeddings, semantic search, model-based disambiguation, and structured outputs.
- Design, tune, and operate keyword and vector retrieval systems, balancing recall, relevance, latency, and scalability across large biomedical datasets.
- Build ranking and relevance layers that determine which concepts and results surface first, incorporating domain-aware weighting and graceful handling of gaps in curated coverage.
- Develop user-facing experiences in Next.js, React, and TypeScript, including question and confirmation flows, result presentation, and status interfaces for long-running processing pipelines.
- Deploy and maintain services across on-premises and high-performance computing Kubernetes environments, working with Helm, StatefulSets, secrets, ingress, GPU scheduling, and scheduled workloads.
- Partner with infrastructure and operations teams to ensure reliable deployments without creating parallel cloud infrastructure.
- Establish rigorous evaluation frameworks for retrieval and concept-mapping quality, using regression suites and measurable benchmarks to detect degradation and validate improvements.
- Implement observability that captures request identifiers, latency, errors, selected concepts, and other decision signals so AI-assisted results can be reviewed and understood.
- Collaborate with researchers, clinicians, and patient-focused stakeholders to understand their needs and translate them into data models, retrieval behavior, and intuitive interfaces.
- Contribute to manuscripts, conference abstracts, posters, and other research outputs, receiving authorship for meaningful contributions.
Requirements
- Bachelor’s degree in Data Science, Computer Science, Bioinformatics, Biomedical Informatics, or a related discipline, or equivalent professional experience; an advanced degree is preferred.
- At least 5 years of experience building and operating production software or data systems, including at least 2 years shipping LLM-powered applications such as agents, retrieval systems, or evaluation frameworks.
- Strong end-to-end retrieval engineering experience covering indexing, query construction, retrieval optimization, and measurement against real-world data.
- Experience evaluating AI or retrieval systems where there is no single correct answer, using golden datasets, offline regression suites, Recall@K, MRR, or comparable metrics.
- Hands-on experience with structured model outputs, tool/function calling, typed schemas, and validation of model-generated data.
- Able to own services across their lifecycle, from schema and data design through deployment, monitoring, troubleshooting, and operation.
- Strong Python skills, including experience with FastAPI, Pydantic, and pytest.
- Deep PostgreSQL experience, including vector search with pgvector or equivalent, full-text search, embedding pipelines, indexing, and query optimization.
- Practical expertise in LLM application engineering, including provider APIs and gateways, prompt and context design, structured generation, and tool use.
- Experience building repeatable data ingestion and transformation pipelines with reliable refresh processes.
- Working knowledge of containers and Kubernetes sufficient to deploy, debug, and operate services in infrastructure you do not directly administer.
- Experience collaborating through Git-based workflows and CI/CD in shared codebases.
- Ability to obtain and maintain a Public Trust Security clearance.
Preferred Experience
- Biomedical ontologies and controlled vocabularies such as MONDO, HPO, UMLS, MeSH, or OBO Foundry resources.
- Familiarity with entity linking, concept normalization, ontology alignment, or embedding-based domain grounding.
- Experience with Helm and on-premises or HPC Kubernetes deployments.
- Exposure to serving open-weight models using technologies such as Ollama or vLLM, model gateways such as LiteLLM, or biomedical embedding models such as MedCPT.
- Background in rare disease, clinical genetics, translational research, NIH programs, or contributions to biomedical standards and open resources.
- Published or presented technical or research work, including papers, conference talks, preprints, technical articles, or open-source contributions.
Benefits
- Competitive salary: Anticipated base compensation of $130,000–$150,000 USD, with actual compensation based on experience, qualifications, skills, and location.
- Medical coverage: 100% medical, dental, and vision coverage for employees.
- Retail savings: 401(k) plan with employer matching of up to 5%.
- Paid time off: PTO and paid holidays.
- Professional development: Educational benefits designed to support career growth.
- Flexible spending accounts: Access to healthcare FSA, parking reimbursement, dependent care assistance, and transportation reimbursement programs.
- Employee referral bonus: Additional rewards for successful employee referrals.
- Remote work: Fully remote position available within the United States.
- Research impact: Opportunity to contribute to meaningful rare disease research alongside scientific and clinical experts.
- Publication opportunities: Contributions to manuscripts, conference abstracts, posters, and other research outputs.