Senior Data Engineer
Capital Technology Group provides expert consulting services in software development, digital transformation, human-centered design, data analytics and visualization, and cybersecurity. Our multidisciplinary teams use agile methodologies to rapidly and incrementally deliver value in close collaboration with federal and commercial clients. For over a decade, we have been trusted to solve complex, mission-critical business challenges.
About the Role
CTG is seeking a Senior Data Engineer to design, build, and maintain scalable, efficient data pipelines and systems following modern data engineering best practices. The Senior Data Engineer will partner with other Data Engineers to evaluate and prototype new tools and technologies, assessing their risks and benefits to deliver exceptional value to our clients.
- Design, build, and maintain scalable data pipelines, ETL/ELT workflows, and data models using Python, Apache Spark (PySpark), Databricks, dbt, SQL (PostgreSQL), and AWS Glue.
- Develop and optimize AWS-native data platforms leveraging AWS Glue, Amazon EMR, Amazon MWAA (Apache Airflow), Lambda, Step Functions, Amazon S3, Redshift, RDS, DMS, and CloudWatch.
- Build high-performance ingestion, transformation, and orchestration workflows for structured and semi-structured data using Apache Iceberg, Parquet, ORC, and Avro.
- Design and optimize analytical data platforms using Amazon Athena, Trino, Hive, OpenSearch, and enterprise data catalog technologies.
- Integrate enterprise and external data sources across relational and NoSQL platforms including PostgreSQL, Oracle, Redshift, GraphDB, and other NoSQL databases.
- Build AI-enabled data solutions using Amazon Bedrock, RAG pipelines, and vector search technologies including Amazon S3 Vector and OpenSearch vector indexes.
- Develop cloud infrastructure using CloudFormation (Infrastructure as Code), GitHub, Harness, and enterprise CI/CD pipelines while leveraging SNS, SQS, and EventBridge for event-driven architectures.
- Improve the reliability, scalability, performance, and maintainability of enterprise data platforms through monitoring, troubleshooting, automation, and continuous optimization.
- Support mission-critical analytics and reporting solutions within large-scale AWS-based federal data environments, implementing solutions that comply with FedRAMP and NIST 800-53 security controls.
- Lead modernization initiatives migrating legacy platforms including IBM DataStage, Hadoop, RunDeck, and shell-based workflows to cloud-native AWS services.
- Mentor junior engineers through technical guidance, architecture discussions, and code reviews while promoting engineering best practices.
- Collaborate with cross-functional teams in an Agile environment to define requirements, deliver high-quality data solutions, and communicate technical concepts effectively to technical and non-technical stakeholders.
Requirements
Applicants MUST BE US Citizens and be able to obtain Public Trust clearance.
- Bachelor's degree in Computer Science, Engineering, or a related technical field.
- 9+ years of professional experience in data engineering or related domains.
- Strong hands-on experience with:
- Databricks, Apache Spark (PySpark), Python, SQL (PostgreSQL), and dbt for large-scale data engineering, ETL/ELT development, data transformation, and data modeling.
- Designing, building, and maintaining AWS-native data platforms using AWS Glue, Amazon EMR, Amazon MWAA (Apache Airflow), AWS Lambda, AWS Step Functions, Amazon S3, Amazon Redshift, Amazon RDS, AWS DMS, and Amazon CloudWatch.
- Developing scalable data pipelines, workflow orchestration, and data integration solutions across enterprise environments.
- Modern data lake technologies including Apache Iceberg and data formats such as Parquet, ORC, and Avro.
- Designing and optimizing solutions using relational and NoSQL databases including PostgreSQL, Redshift, Oracle, GraphDB, and other NoSQL platforms.
- Building reliable, high-performance data platforms through performance tuning, system optimization, and enterprise-scale ETL/ELT architectures.
- Java development and modern CI/CD practices using Harness.
- Strong analytical and problem-solving skills.
- Experience working in agile, iterative software development environments.
- Ability to quickly learn and apply new technologies and domain knowledge.
- Excellent written and verbal communication skills, with the ability to explain complex topics to diverse audiences.
Skills
Who You Are:
- A strategic data engineer who enjoys designing complex systems and solving complex challenges.
- Strong in modern cloud-based solution design.
- Comfortable balancing business needs with technical constraints and long-term strategy.
- A strong communicator.
- Collaborative, proactive, and comfortable navigating ambiguity.
Nice to Have:
- Experience supporting analytics, data engineering, or modernization initiatives for financial regulators, capital markets, or other highly regulated environments is a plus.
- Experience with Kafka (streaming/data pipelines).
- Experience with Docker and Kubernetes for containerization and orchestration.
- Proficiency with Splunk for log aggregation and system monitoring.
- Experience using Terraform for infrastructure automation and management.
- Strong SQL skills, including performance tuning and complex query design.
Benefits
- Remote Work (Hybrid roles will be specified in the job post).
- Competitive Compensation Package.
- Medical, Dental, and Vision.
- Life Insurance, Short/Long Term Disability.
- Employee Assistance Program.
- 401(k) with 4% matching.
- Liberal PTO vacation policy.
- Generous Annual Continuing Education.
- Annual Wellness Budget.
- Bonus Incentive Programs (Employee referrals and performance-based rewards).
Pay
We are committed to offering a competitive salary for this position, with an estimated range of $130k to $165k annually. Please note that this range is intended to provide a general idea of what to expect; however, the final offer may vary based on experience, skills, and other factors.