Junior Data Engineer (DEA)
hatch I.T. · Arlington, VA · 3 days ago
On-siteInformation TechnologyFull-time
About The Role
Expression is seeking a Junior Data Engineer to support the Drug Enforcement Administration (DEA) Investigative Case and Data Ecosystem (ICDE) modernization effort. The program is focused on replacing fragmented legacy investigative and case management capabilities with a centralized, secure, scalable, and integrated ecosystem.
Responsibilities
- Assist in designing, developing, and maintaining ETL pipelines for ingesting, transforming, and loading datasets under the guidance of more experienced engineers.
- Migrate legacy data into the modernized DEA investigative and case management ecosystem.
- Analyze data structures, source-to-target mappings, and data-quality checks to identify issues or gaps.
- Ensure data accuracy, consistency, and integrity by executing validation queries, profiling datasets, and supporting data-cleaning efforts.
- Contribute to data migration activities, including mapping source data to target systems and executing migration scripts.
- Support data cleaning, standardization, classification, and tagging activities.
- Collaborate with analysts, data scientists, engineers, and business stakeholders to translate requirements into data transformations and pipeline updates.
- Monitor pipeline performance and assist with troubleshooting operational issues, escalating complex issues as appropriate.
- Participate in validation activities and help document migration results and identified data-quality issues.
- Participate in Agile ceremonies and support data engineering activities throughout development sprints.
- Document data flows, transformation logic, mappings, and operational processes to support maintainability and knowledge sharing.
Qualifications
- Bachelor's degree in STEM fields and 0–1 years of professional experience.
- Foundational exposure to data analysis, ETL concepts, or data migration through coursework, internships, personal projects, or early professional experience.
- Working knowledge of SQL, including basic queries, joins, and aggregations.
- Introductory-level experience with Python for data manipulation or automation tasks.
- Familiarity with ETL or workflow tools such as Apache Airflow, Talend, or similar technologies.
- Understanding of basic data warehousing concepts, including staging, fact/dimension models, or schema structures.
- Exposure to cloud-based data storage or compute platforms such as AWS S3/Redshift, Google BigQuery, Azure Storage, or similar environments.
- Foundational understanding of data profiling, validation, and data quality.
- Strong communication and documentation skills.
- Ability and willingness to learn in a fast-paced, evolving technical environment.
Preferred Qualifications
- Hands-on or coursework experience with AWS, Azure, or Google Cloud Platform.
- Exposure to Hadoop, Spark, or other distributed-processing technologies.
- Familiarity with Power BI, Tableau, Looker, or similar visualization tools.
- Familiarity with Git or similar version-control systems.
- Experience building or supporting automated data workflows using orchestration tools or scheduled scripting.
- Exposure to Agile development methodologies.
- Coursework, internship, project, or professional exposure to data migration or system-modernization activities.