Senior Databricks Engineer
CGI · United States · Today
RemoteRemoteEngineering$101k/yrContract
About the role
The Senior Databricks Engineer will design, develop, test, and deploy scalable batch and real-time data engineering solutions while reverse-engineering existing AWS-based data processing to identify transformations, business rules, dependencies, and integration requirements.
Responsibilities
- Design, develop, test, and deploy scalable Databricks data pipelines and transformation workflows as part of a greenfield Databricks implementation.
- Analyze and reverse engineer existing AWS-based data processing solutions, including undocumented code and pipeline logic, to identify data transformations, business rules, dependencies, orchestration, and integration requirements that must be preserved or redesigned within Databricks.
- Build, enhance, and maintain Bronze, Silver, and Gold data layers supporting customer data models and downstream consumption requirements.
- Design and implement scalable ingestion solutions using Databricks Auto Loader, including schema management, checkpointing, incremental processing, backfill/reprocessing, and production-scale file ingestion patterns.
- Develop and operate Lakeflow Declarative Pipelines (formerly Delta Live Tables) supporting batch and streaming workloads, data quality expectations, quarantine/error handling, pipeline dependencies, monitoring, and recovery.
- Develop real-time and batch data processing solutions using Databricks, Structured Streaming, and associated technologies.
- Implement robust transformation logic using Apache Spark, PySpark, Spark SQL, and Delta Lake.
- Design and optimize Delta Lake solutions using appropriate MERGE patterns, Change Data Feed, partitioning/clustering, retention, and optimization strategies.
- Design solutions that support reliable ingestion, transformation, enrichment, and delivery of high-volume customer data.
- Design and implement appropriate Unity Catalog governance structures, including catalogs, schemas, volumes, external locations, storage credentials, access controls, lineage, and promotion across development, test, and production environments.
- Develop and support integrations between Databricks, Adobe, AWS services, and other identified downstream systems.
- Analyze existing AWS Glue, Amazon S3, Amazon Redshift, Redshift Spectrum, and supporting AWS services to determine existing processing behavior, embedded business rules, dependencies, security requirements, and target state Databricks implementation.
- Establish reusable Databricks Workflows, deployment as code patterns, frameworks, utilities, and engineering standards appropriate for a greenfield implementation and subsequent development teams.
- Apply Databricks engineering best practices for data quality, performance optimization, scalability, observability, reliability, security, and maintainability.
- Perform performance tuning and optimization of Spark workloads, pipelines, queries, streaming processes, and data structures, including analysis of partitioning, shuffle behavior, join strategies, data skew, and Spark execution characteristics.
- Troubleshoot complex data pipeline, integration, data quality, performance, and production issues and participate in root cause analysis and remediation.
- Develop automated validation and testing approaches to ensure the accuracy, completeness, and reliability of data products.
- Evaluate legacy processing logic to distinguish required business functionality from source platform workarounds, avoiding unnecessary replication of legacy technical constraints within the target architecture.
- Participate in design discussions, peer code reviews, technical reviews, and solution refinement.
- Work within client-established development, CI/CD, security, governance, and deployment processes.
- Produce and maintain technical documentation, including data flows, implementation details, development patterns, operational procedures, reverse engineering findings, and troubleshooting guidance.
- Provide hands-on knowledge transfer and mentoring to client engineers to strengthen internal Databricks engineering capabilities and enable long-term platform ownership.
- Leverage approved AI-assisted development tools, where appropriate, to improve development velocity, code quality, testing, documentation, and engineering efficiency.
- Identify and recommend opportunities for automation and engineering improvement that increase delivery speed, solution quality, operational reliability, and recovery from production issues.
Requirements
- 6+ years of data engineering or software engineering experience, including significant experience designing, developing, and supporting enterprise-scale data platforms.
- 3+ years of hands-on Databricks experience developing and operating production data engineering solutions.
- Strong hands-on experience with Databricks Auto Loader, including incremental cloud object storage ingestion, schema inference and evolution, schema hints, rescued data handling, checkpoint/state management, backfill and reprocessing strategies, and scalable file discovery approaches.
- Strong hands-on experience with Lakeflow Declarative Pipelines (formerly Delta Live Tables), including batch and streaming pipelines, data quality expectations, failed record/quarantine handling, streaming tables and materialized views, incremental versus full refresh processing, dependency management, monitoring, and alerting.
- Strong hands-on experience with Unity Catalog in a production environment, including catalog/schema/volume design, external locations, storage credentials, grants, row and column-level access controls, lineage, and asset promotion across environments.
- Strong hands-on experience with Delta Lake, including MERGE patterns, Change Data Feed, time travel, OPTIMIZE, Z ORDER and/or liquid clustering, VACUUM and retention policies, and partitioning strategies at scale.
- Strong hands-on experience with Apache Spark, PySpark, and Spark SQL, including development of complex data transformations and production performance tuning.
- Strong hands-on experience with Structured Streaming, including watermarking, late-arriving data, stateful processing, checkpointing, and exactly once processing considerations.
- Strong hands-on experience with Medallion Architecture, including personally building Bronze, Silver, and Gold data layers and making appropriate decisions regarding responsibilities and boundaries between layers.
- Experience with Databricks Workflows and deployment as code, including job dependencies, retry/failure handling, alerting, and Git-based environment promotion using Databricks Asset Bundles, Terraform, or comparable automation.
- Strong understanding of distributed data processing, data optimization, scalability, and production pipeline resiliency.
- Experience implementing production-grade approaches for data quality, error handling, monitoring, logging, observability, and pipeline recovery.
- Experience with CI/CD, automated deployment practices, Git-based source control, branching, pull requests, and peer code reviews.
- AWS Source Environment Experience: Demonstrated experience with AWS Glue ETL jobs using PySpark and Python, including DynamicFrames versus DataFrames, job bookmarks, connections, job parameters, worker sizing, and Glue runtime/version considerations.
- Strong understanding of Amazon S3, including bucket/prefix structures, partitioning, file formats and compression, access patterns, and considerations affecting downstream data ingestion.
- Experience with Amazon Redshift, including data structures, distribution and sort strategies, COPY/UNLOAD patterns, stored procedures, and the ability to identify transformation or business logic embedded within the warehouse.
- Strong understanding of AWS IAM and data access patterns, including roles, assumed role relationships, policies, and authentication mechanisms sufficient to trace how existing data pipelines access source and target data and translate those requirements into appropriate Databricks and Unity Catalog access patterns.
- Working knowledge of AWS Glue Crawlers and the Glue Data Catalog, including schema inference, schema drift, partition management, and the relationship between existing catalog definitions and the target Unity Catalog model.
- Working knowledge of Redshift Spectrum and external S3-backed data access patterns, with the ability to understand how existing external schemas and datasets should be represented within Databricks and Unity Catalog.
- Working knowledge of supporting AWS data, orchestration, monitoring, and security services such as Lambda, Step Functions, EventBridge, Athena, CloudWatch, and Secrets Manager.
- Demonstrated ability to trace an end-to-end AWS data pipeline across multiple services, identifying data sources, transformations, orchestration, dependencies, security/access requirements, failure handling, and downstream consumers necessary to support migration to Databricks.
- Reverse Engineering and Migration Experience: Demonstrated experience reverse engineering undocumented legacy data pipelines where business rules and processing requirements are embedded within application, ETL, orchestration, or database code.
- Demonstrated experience performing data reconciliation and parity validation between legacy and rebuilt pipelines, investigating discrepancies and confirming equivalent business outcomes.
- Ability to distinguish business logic that must be preserved from technical workarounds or constraints of the legacy platform, and redesign appropriately for a modern Databricks architecture.
- Experience working in Agile delivery environments and delivering against prioritized product backlogs.
- Strong communication and collaboration skills with the ability to work effectively across engineering, architecture, product, and business teams.