ML Data Engineer
About The Role
We are seeking a highly skilled Machine Learning Data Engineer to join our dynamic team remotely. In this role, you will be responsible for designing, building, and maintaining large-scale data systems that support AI training and evaluation pipelines. The ideal candidate will possess extensive experience in data engineering, particularly supporting machine learning and AI workloads, with a focus on ingestion, transformation, quality assurance, lineage, and high-throughput data delivery. Your work will directly impact the efficiency and effectiveness of our AI models by ensuring that data infrastructure is robust, scalable, and optimized for performance.
Qualifications
- Bachelor’s or Master’s degree in Computer Science, Data Science, or a related field
- Six or more years of experience in data engineering, with a focus on supporting machine learning or AI workloads
- Proficiency in Python and at least one JVM or systems programming language
- Deep experience with modern data processing frameworks such as Spark, Ray, or Beam
- Hands-on experience managing petabyte-scale storage and data pipelines
- Strong understanding of distributed systems, data modeling, and storage formats
- Experience with dataset versioning, lineage, and reproducibility in ML workflows
- Knowledge of high-throughput data loading techniques optimized for GPU/accelerator training
- Solid software engineering practices including testing, CI/CD, and code review
- Excellent communication skills and ability to collaborate across teams
Responsibilities
- Design and operate large-scale data pipelines to support AI training, evaluation, and continuous improvement processes
- Develop ingestion systems capable of handling diverse modalities such as text, images, audio, video, and structured data
- Implement data cleaning, deduplication, filtering, and quality assurance processes at petabyte scale
- Create dataset versioning, lineage, and provenance tracking systems to ensure reproducibility
- Build high-throughput data loading systems that optimize GPU utilization during training sessions
- Develop labeling workflows, active learning pipelines, and human-in-the-loop data enhancement systems
- Construct evaluation datasets with strict controls for integrity and contamination prevention
- Apply data privacy, redaction, and consent enforcement measures throughout data pipelines
- Collaborate with ML researchers and engineers to align data infrastructure with model development needs
- Monitor data quality, drift, and pipeline health to ensure consistent data delivery and reliability
- Optimize data infrastructure performance through compression, caching, and format selection
- Document data systems, schemas, and operational procedures for internal teams
- Stay updated with the latest research, tools, and best practices in AI data infrastructure
Benefits
- Competitive salary range of $100,000 to $150,000 annually
- Remote work flexibility within the U.S.
- Full-time employment with direct W2 status
- Opportunities for professional growth and career advancement
- Collaborative and innovative work environment
- Access to cutting-edge projects and technologies in AI and data engineering
- Comprehensive benefits package (details to be provided during onboarding)
Equal Opportunity
Bright Vision Technologies (BV Teck) is committed to fostering an inclusive and diverse work environment. We provide equal employment opportunities to all qualified applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, veteran status, or any other protected characteristic under applicable federal, state, or local laws. We prohibit discrimination and harassment of any kind and are dedicated to ensuring a workplace where everyone can thrive and contribute to our collective success.