Data Scientist
TheCorporate · Woodlawn, MD · Yesterday
On-siteEngineeringContract
Job Summary We are seeking an experienced Data Scientist with strong expertise in Python, SQL, Natural Language Processing (NLP), entity resolution, record linkage, and large-scale data processing. The ideal candidate will have hands-on experience developing sophisticated data processing and entity resolution pipelines, working with structured and unstructured data, optimizing complex database operations, and translating advanced matching and scoring algorithms into clear business logic. This role will support the end-to-end delivery of data solutions, including data discovery, cleansing, validation, testing, deployment, monitoring, and optimization. Key Responsibilities Design, develop, implement, and maintain advanced data processing and entity resolution pipelines using Python and SQL.Apply NLP and information extraction techniques to process and extract insights from unstructured data, including:Named Entity Recognition (NER)String distance metricsTF-IDFCosine similarityPhonetic encodingAddress standardizationClean, transform, standardize, and manage large-scale datasets from complex and diverse data sources.Ensure data integrity, quality, security, and privacy throughout data processing workflows.Develop and optimize complex SQL queries and database operations for performance, scalability, and reliability.Develop solutions for record linkage, entity resolution, and data deduplication.Implement blocking and indexing strategies to efficiently process large datasets.Support end-to-end data solution delivery, including:Data discoveryData validationTestingDeploymentPost-implementation monitoringParticipate in code reviews and enforce version control, coding standards, reproducibility, and data privacy practices.Translate complex algorithmic decisions, including match thresholds and probabilistic scoring, into clear and understandable business logic.Communicate technical findings and algorithmic approaches effectively to both technical teams and non-technical stakeholders. Qualifications & Requirements 10+ years of IT industry experience.Bachelor's degree in Statistics, Applied Mathematics, Computer Science, Information Science, or a related field (or equivalent experience).Proven expertise in NLP and information extraction, including:Named Entity Recognition (NER)String distance metrics, TF-IDF, Cosine similarity, Phonetic encoding, Address standardizationStrong hands-on Python skills for building analytics solutions and scalable data pipelines.Advanced SQL skills, including query optimization and complex database operations.Proficiency with Regex for text processing, pattern matching, and data cleansing.Experience with record linkage and deduplication libraries such as:Splink / FastLinkDedupeRecordlinkageExperience using spaCy for NLP and scikit-learn for machine learning and clustering.Strong understanding of data quality, integrity, security, and privacy principles.Excellent written and verbal communication skills, with the ability to translate technical concepts for diverse audiences. Preferred / Nice-to-Have Qualifications Experience delivering data initiatives in federal, state, or local government environments.Experience retrieving, processing, and migrating data from legacy and distributed systems, including:PostgreSQL, DB2, Oracle, SQL Server, Hadoop, Flat filesExperience with Jenkins and CI/CD for automating data validation and processing pipelines.Experience with pipeline automation, scheduling, and monitoring tools for complex data cleansing and processing jobs.Proven ability to own data pipeline architecture end-to-end, from initial data discovery through post-implementation monitoring.Experience designing scalable data processing solutions for large and complex datasets.Experience developing reproducible and production-ready data science solutions. Technical Skills Programming: Python, SQL, RegexNLP / Information Extraction: NER, spaCy, TF-IDF, Cosine Similarity, String Distance Metrics, Phonetic Encoding, Address StandardizationMachine Learning: scikit-learn, ClusteringEntity Resolution: Record Linkage, Entity Resolution, Deduplication, Blocking, Indexing, Splink/FastLink, Dedupe, RecordlinkageDatabases: PostgreSQL, DB2, Oracle, SQL Server, Greenplum/Hadoop environments as applicableData Engineering: Data Pipelines, Data Cleansing, Data Transformation, Data Validation, Data Migration, Pipeline MonitoringDevOps / CI/CD: Jenkins, Version Control, Automated Data ValidationData Concepts: Data Integrity, Data Quality, Data Privacy, Data Security, Reproducibility Education Bachelor's degree in Statistics, Applied Mathematics, Computer Science, Information Science, or a related field — or equivalent professional experience. [Add information regarding benefits and perks here] Skills: sql,data,nlp,data validation,public trust,pipelines,regex,processing