Data Engineer
Capital Technology Group · Silver Spring, MD · 1 mo ago
RemoteRemoteInformation Technology$110k–$140k/yrFull-time
About the role
CTG is seeking a Data Engineer to design, build, and maintain scalable, efficient data pipelines and systems following modern data engineering best practices.
Responsibilities
- Design, build, and maintain scalable data pipelines, ETL/ELT workflows, and data models using Python, Apache Spark (PySpark), Databricks, dbt, SQL (PostgreSQL), and AWS Glue.
- Develop and optimize AWS-native data platforms leveraging AWS Glue, Amazon EMR, Amazon MWAA (Apache Airflow), Lambda, Step Functions, Amazon S3, Redshift, RDS, DMS, and CloudWatch.
- Build high-performance ingestion, transformation, and orchestration workflows for structured and semi-structured data using Apache Iceberg, Parquet, ORC, and Avro.
- Design and optimize analytical data platforms using Amazon Athena, Trino, Hive, OpenSearch, and enterprise data catalog technologies.
- Integrate enterprise and external data sources across relational and NoSQL platforms including PostgreSQL, Oracle, Redshift, GraphDB, and other NoSQL databases.
- Develop AI-enabled data solutions using Amazon Bedrock, RAG pipelines, and vector search technologies including Amazon S3 Vector and OpenSearch vector indexes.
- Develop cloud infrastructure using CloudFormation (Infrastructure as Code), GitHub, Harness, and enterprise CI/CD pipelines while leveraging SNS, SQS, and EventBridge for event-driven architectures.
- Improve the reliability, scalability, performance, and maintainability of enterprise data platforms through monitoring, troubleshooting, automation, and continuous optimization.
- Migrate legacy platforms including IBM DataStage, Hadoop, RunDeck, and shell-based workflows to cloud-native AWS services.
- Lead modernization initiatives.
- Mentor junior engineers through technical guidance, architecture discussions, and code reviews while promoting engineering best practices.
- Collaborate with cross-functional teams in an Agile environment to define requirements, deliver high-quality data solutions, and communicate technical concepts effectively to technical and non-technical stakeholders.
Qualifications
- Bachelor's degree in Computer Science, Engineering, or a related technical field.
- 4+ years of professional experience in data engineering or related domains.
- Strong hands-on experience with:
- Databricks, Apache Spark (PySpark), Python, SQL (PostgreSQL), and dbt for large-scale data engineering, ETL/ELT development, data transformation, and data modeling.
- Designing, building, and maintaining AWS-native data platforms using AWS Glue, Amazon EMR, Amazon MWAA (Apache Airflow), AWS Lambda, AWS Step Functions, Amazon S3, Amazon Redshift, Amazon RDS, AWS DMS, and Amazon CloudWatch.
- Developing scalable data pipelines, workflow orchestration, and data integration solutions across enterprise environments.
- Working with modern data lake technologies including Apache Iceberg and data formats such as Parquet, ORC, and Avro.
- Designing and optimizing solutions using relational and NoSQL databases including PostgreSQL, Redshift, Oracle, GraphDB, and other NoSQL platforms.
- Building reliable, high-performance data platforms through performance tuning, system optimization, and enterprise-scale ETL/ELT architectures.
- Java development and modern CI/CD practices using Harness.
- Strong analytical and problem-solving skills.
- Experience working in agile, iterative software development environments.
- Ability to quickly learn and apply new technologies and domain knowledge.
- Excellent written and verbal communication skills, with the ability to explain complex topics to diverse audiences.
Qualifications (Nice to Have)
- Experience supporting analytics, data engineering, or modernization initiatives for financial regulators, capital markets, or other highly regulated environments.
- Experience with Kafka (streaming/data pipelines).
- Experience with Docker and Kubernetes for containerization and orchestration.
- Proficiency with Splunk for log aggregation and system monitoring.
- Experience using Terraform for infrastructure automation and management.
- Strong SQL skills, including performance tuning and complex query design.
Pay
We are committed to offering a competitive salary for this position, with an estimated range of $110k to $140k annually. Please note that this range is intended to provide a general idea of what to expect; however, the final offer may vary based on experience, skills, and other factors. The stated range is not a guarantee and is subject to change.