Data Science Engineer Internship
CCC Intelligent Solutions Inc. (CCC) is a leading cloud platform for the multi-trillion-dollar insurance economy, creating intelligent experiences for insurers, repairers, automakers, part suppliers, and more. At CCC, we’re making life just work by empowering more than 35,000 businesses with industry-leading technology to get drivers back on the road and to health quickly and seamlessly. We’re pushing boundaries with innovative AI solutions that simplify and enhance the claims and repair journey.
About the role
Our program is designed to #CCCJumpstart your career! At CCC, you will work and learn alongside innovative and inspiring leaders and gain valuable technical experience while working on real business solutions in a corporate setting. You'll join the Data Engineering team that builds the data and AI foundation for the company: the curated, governed, well-modeled data that powers analytics, machine learning, and generative AI across CCC.
Responsibilities
- Build end-to-end pipelines that turn raw data into fully curated, enhanced datasets ready for analytics and model training.
- Work with data across all three shapes: structured (relational databases, transactional tables), semi-structured (JSON, XML, compressed/zipped files), and unstructured (claim photos, estimate PDFs, repair notes, audio, and video).
- Produce reusable data building blocks, data models, and data flows for varying client demands — dimensional models, data feeds, dashboard reporting, and data science research and exploration.
- Develop and optimize Spark jobs on Databricks, including Delta Lake tables and incremental and batch processing patterns.
- Apply data governance practices using Unity Catalog and AWS Glue Data Catalog: cataloging and lineage, access controls, PII handling, schema management, and data quality validation.
- Use Apache Airflow to build DAGs that handle cross-DAG dependencies, orchestrate concurrent pipelines, monitor tasks, and alert when SLAs are not met.
- Partner with data scientists and ML engineers to make datasets model-ready, including preparing multimodal data for AI and computer vision use cases.
Technical environment
- Python and Spark
- Databricks, Delta Lake, and Unity Catalog
- AWS ecosystem (S3, EMR, Glue, Athena)
- Airflow for scheduling and monitoring of big data ETL pipelines
- SQL for data profiling and validation
- Unix commands and scripting
- Distributed computing and lakehouse fundamentals
- Generative AI and Agentic AI fundamentals
Requirements
- In pursuit of an Associate, Bachelor's or Master's degree throughout your internship.
- Strong collaborative skills and ability to work well with a team.
- A strong interest in computer science and/or related fields.
- Knowledge of Python, Machine Learning, or Statistics/Modeling through coursework or project work.
Pay
The hourly range is $20.00 - $43.00 per hour. Pay is based on factors including school year, program of study, and role responsibilities.
Benefits
- 401K Match
- Paid time off
- Annual Incentive Plan
- Performance Bonus
- Comprehensive health insurance
- Adoption Assistance
- Tuition Reimbursement
- Wellness Programs
- Stock Purchase Plan options
- Employee Resource Groups