Urgent Need Senior Data Engineer – Financial Fraud Analytics
Vinsys Information Technology Inc · Washington, DC · 1 mo ago
Information TechnologyFull-time
Role
The Senior Data Engineer will design, implement, maintain, and improve an integrated and flexible data architecture within SBA OIG's Microsoft Azure environment.
Responsibilities
- Provide authoritative expertise in data-engineering methods and best practices.
- Apply code-first development approaches and modern pipeline-design patterns.
- Design and maintain a secure, stable, scalable, and flexible data architecture.
- Manage data assets through source control.
- Design, implement, and maintain ELT/ETL pipelines.
- Develop pipelines using Azure Synapse and Azure Machine Learning.
- Migrate source data into Azure Data Lake Storage.
- Review, maintain, and improve existing architecture and pipelines.
- Conduct periodic reviews to identify bottlenecks, deprecated dependencies, and architecture drift.
- Implement pipeline quality controls, error handling, logging, monitoring, and validation checks.
- Incorporate source control into data pipelines and analytics codebases.
- Optimize data ingestion, processing, storage, and retrieval.
- Work with structured, semi-structured, and unstructured data.
- Use modern columnar formats, including Parquet.
- Normalize common entity attributes, including names, addresses, telephone numbers, and other identifying information.
- Develop self-service capabilities that allow SBA OIG analysts to query and export data.
- Cook with data scientists to support machine-learning models and analytical pipelines.
- Develop SOPs for authoring, developing, validating, publishing, executing, and monitoring pipelines and assets.
- Develop data dictionaries, entity-relationship diagrams, pipeline maps, and architecture documentation.
- Expand the environment with additional datasets and services as requested.
- Establish intake, testing, and production-deployment procedures.
- Monitor pipelines to ensure performance and regular dataset updates.
- Recommend architecture changes that reduce cloud costs.
- Evaluate emerging AI, automation, coding-assistant, and LLM-assisted data-engineering capabilities.
Required Qualifications
- Candidates Must Possess One Of The Following Bachelor's degree in Data Engineering, Computer Science, Data Science, Machine Learning, Mathematics, or a related field; or Five years of applied work experience in one or more of these fields.
- Candidates must have at least five years of hands-on experience in each of the following: Maintaining SQL databases, Conducting advanced SQL and T-SQL operations, Designing, implementing, and maintaining ELT/ETL processes in cloud-based data-analytics environments.
- Candidates must have at least three years of hands-on experience in each of the following: Working with Azure Synapse, Working with Azure Machine Learning, Working with modern data-stack technologies, Manipulating data using Python, Using Pandas.
Preferred Qualifications
- Microsoft DP-203 certification or equivalent.
- PySpark or Polars experience.
- Experience developing reusable and modular code.
- Experience implementing pipelines and infrastructure using Python SDKs, command-line tools, REST APIs, or Infrastructure-as-Code tools.
- Experience implementing source-control and CI/CD workflows.
- Familiarity with AI coding assistants and LLM integration patterns.