Founding Data Engineer
Glocomms · San Francisco County, CA · 1 mo ago
ScienceFull-time
About the role
A high-impact Data Engineer role focused on building the core data layer for AI systems that improve decision-making in drug development. The team is working to unify fragmented experimental data, research, and manual judgment into a data-driven approach for predicting safety risks earlier in the process. You’ll design and scale pipelines, build internal APIs and tooling, and transform messy raw data into clean, structured datasets for modeling and analysis. The role includes automating data ingestion, improving data quality, and supporting model-related workflows, with exposure to LLM-driven workflows for extracting and structuring information.
Responsibilities
- Own and scale core data infrastructure across research, ML, and product systems
- Build and maintain data pipelines for ingesting, processing, validating, and serving large, complex datasets from multiple sources
- Develop internal platforms that connect experimental workflows, data capture, processing, and downstream usage
- Transform raw, messy data into clean, versioned, and ML-ready datasets
- Design and build APIs and data tools that make it easy for different teams to access and use data
- Work on systems that automate data ingestion, cleaning, normalization, and structuring across a variety of inputs
- Support and scale model-related workflows, including batch processing and inference pipelines
- Implement data quality systems (validation, testing, monitoring, lineage, observability) to ensure reliability
- Partner closely with domain experts to understand workflows and translate them into scalable infrastructure
- Help support internal and external data delivery, including datasets, outputs, and derived insights
- Build systems that improve speed and efficiency across the organization
Requirements
- Experience building and maintaining large-scale data platforms used by multiple teams
- Comfortable working with messy, unstructured, or heterogeneous datasets
- Ability to operate across backend engineering, data systems, and infrastructure
- Strong focus on data quality, correctness, and reproducibility
- Experience working with cross-functional teams and translating real-world needs into technical solutions
- Familiarity with AI/ML or LLM-related data workflows is a plus
- Comfortable operating in ambiguous, fast-moving environments
- Interested in owning critical systems and having broad impact
- Curiosity across technical and applied problem spaces
Skills
- Strong Python and SQL fundamentals
- Experience with distributed systems or large-scale data processing frameworks
- Cloud infrastructure (any major provider)
- Infrastructure as code and modern deployment practices
- Data platforms (warehouses, lakes, or similar storage systems)
- Workflow orchestration tools
- Experience building APIs or internal data tooling
- Exposure to ML or LLM-related infrastructure
- Experience handling large-scale datasets
What tends to work well in this environment
- Moves quickly and takes ownership
- Strong engineering judgment and attention to detail
- Proactively identifies problems and builds solutions
- Balances speed with reliability
- Thinks in systems rather than one-off fixes
- Comfortable navigating complexity and changing requirements
- Enjoys working with technical and research-oriented teams
- Focuses on building scalable, long-term solutions
- Motivated by high-impact work in an early-stage environment