Machine Learning Engineer Intern — AI for Science
About the Role
We’re building an AI agent that helps scientists make sense of large collections of historical experimental data. The first component ingests many heterogeneous data tables and automatically works out how they relate to one another. The role sits at the intersection of machine learning, large language models, and data engineering, in close collaboration with academic domain experts. This internship provides an opportunity to gain hands-on experience applying AI techniques to real-world scientific data challenges, while learning how machine learning, LLMs, and data engineering can work together to build practical AI systems.
Responsibilities
- Exploring and profiling messy, real-world scientific tables (CSV/Excel), extracting structure and metadata.
- Contributing to a relationship-inference pipeline that combines classical data-integration signals (schema and value matching, key detection) with LLM-based semantic reasoning over column meanings and provenance.
- Supporting the development of evaluation methods that measure system performance against expert-provided ground truth.
- Loading, cleaning, and normalizing heterogeneous tabular datasets and building reusable data-profiling tooling.
- Prototyping and iterating on an LLM/agent pipeline that classifies how pairs of tables relate (parallel / shared / hierarchical).
- Designing and maintaining benchmarks and metrics to evaluate system accuracy, and running experiments to improve performance.
- Working with database schemas, including joins, keys, and data modeling, to represent and query discovered relationships.
- Collaborating with mentors and domain scientists to understand requirements and turn feedback into concrete improvements.
- Documenting findings, experiments, and technical approaches throughout the project.
Required Qualifications
- BS/MS (in progress or completed) in Data Science, Computer Science, or a closely related field.
- Strong Python skills, including data-wrangling libraries (e.g., pandas, NumPy).
- Solid grounding in machine learning fundamentals.
- Hands-on database experience, including SQL, schema design, and joins.
Preferred Qualifications
- Experience with LLMs / AI agents (prompting, RAG, tools like LangChain).
- Prior work with scientific or experimental datasets.
- Familiarity with data integration, schema matching, or entity resolution.
- Good software habits such as version control, testing, and clear documentation.
What We Offer
- Fully remote with flexible schedule.
- Collaborate with team members from leading tech firms (including MAMAA).
- Work on high-impact, real-world AI projects.
- Great for students (supports CPT/OPT).
- Opportunity for recommendation letters, referrals, and future growth.
- Direct mentorship from professionals in product, marketing, and AI.
Object Tech, Inc. is a leading technology company providing full-stack AI solutions to reform tech enterprises’ technology R&D and production. We are committed to fostering a dynamic and inclusive work environment where creativity and innovation thrive.