Data Collection/Processing Intern
Juxta · San Francisco, CA · 3 wk ago
On-siteAccountingInternship
About the Role
This is a hands-on role that combines field data collection, quality assurance, and data preprocessing. You’ll work directly with Juxta’s data collectors and accompany them to different collection locations, ensuring the data we collect is accurate, consistent, and ready for model training.
Responsibilities
- Support field data collection: Travel with Juxta’s data collectors to different locations and help oversee collection sessions.
- Ensure collection quality: Verify that collectors follow correct procedures, equipment is properly configured, and collected data meets Juxta’s quality standards.
- Identify issues in real time: Catch problems like incorrect setups, missing data, inconsistent procedures, or corrupted files before they impact a collection session.
- Validate collected data: Review datasets post-collection to confirm files are complete, correctly organized, and usable.
- Preprocess data: Clean, organize, format, filter, and validate raw data for downstream training workflows.
- Document collection sessions: Maintain clear records of collection conditions, issues encountered, and changes affecting the dataset.
- Improve processes: Work with the team to identify recurring data-quality issues and enhance collection and preprocessing workflows.
Requirements
- Strong attention to detail and ability to spot inconsistencies.
- Comfort working with files, datasets, and basic technical workflows.
- Ability to follow and enforce detailed data-collection protocols.
- Strong organizational and problem-solving skills.
- Comfortable working independently and as part of a field team.
- Willingness to travel locally to data collection locations.
- Interest in AI, machine learning, data engineering, robotics, computer vision, or related fields.
- Currently pursuing or recently completed a degree in Computer Science, Data Science, Engineering, Statistics, or a related technical field.
Nice to Have
- Experience working with Python.
- Familiarity with data cleaning/preprocessing tools (e.g., NumPy, Pandas).
- Experience with large datasets, sensor data, images, video, audio, or multimodal data.
- Knowledge of machine learning training pipelines or dataset preparation.
- Previous experience in field research, data collection, QA, or technical operations.
What You’ll Learn
Gain direct exposure to building real-world datasets for AI, from raw data collection to model training preparation. Work closely with Juxta’s technical and operations teams, developing practical skills in data quality, preprocessing, AI pipelines, and large-scale data operations. Ideal for those who enjoy hands-on work, value precision, and want to understand how high-quality training data is created.