Technical Lead - DataBricks
Wipro · Minneapolis, MN · 3 days ago
Engineering$60k–$135k/yrFull-time
Location: Dallas, TX (3 days onsite per week)
About the role
We are looking for a highly skilled and hands-on Databricks Lead Data Engineer to design, build, optimize, and manage scalable data engineering solutions using Databricks, Apache Spark, PySpark, SQL, and cloud-based data platforms. The ideal candidate should have strong technical leadership capabilities along with deep hands-on experience in developing production-grade data pipelines, lakehouse architectures, and enterprise data solutions.
Responsibilities
- Lead the design, development, and implementation of scalable data pipelines and data lakehouse solutions using Databricks, PySpark, Spark SQL, and Delta Lake.
- Work hands-on in building batch and streaming ETL/ELT pipelines from multiple source systems into cloud data platforms.
- Design and implement medallion architecture layers such as Bronze, Silver, and Gold for efficient data processing and analytics consumption.
- Optimize Spark jobs, Databricks notebooks, clusters, workflows, and SQL queries for performance, reliability, and cost efficiency.
- Collaborate with business stakeholders, architects, data analysts, and data scientists to understand requirements and translate them into robust technical solutions.
- Provide technical leadership, code reviews, best practices, and mentoring support to junior and mid-level data engineers.
- Implement data quality checks, data validation rules, monitoring, logging, error handling, and restartability mechanisms.
- Ensure adherence to data governance, security, access control, and compliance standards across data engineering solutions.
- Support production deployments, troubleshoot pipeline failures, perform root cause analysis, and drive continuous improvement.
Requirements
- 5–8 years of experience in data engineering.
- Strong hands-on experience in Databricks development, including notebooks, workflows, jobs, clusters, Delta Lake, and Unity Catalog.
- Advanced programming experience in PySpark, Python, Spark SQL, and SQL.
- Strong understanding of data engineering concepts, data warehousing, data lakehouse architecture, ETL/ELT design, and data modeling.
- Experience in building scalable data pipelines on cloud platforms such as Azure, AWS, or GCP.
- Hands-on experience with orchestration tools such as Azure Data Factory, Airflow, Databricks Workflows, or similar tools.
- Experience working with file formats such as Parquet, Avro, JSON, CSV, and Delta format.
- Good understanding of performance tuning techniques including partitioning, caching, broadcast joins, cluster sizing, compaction, and query optimization.
- Experience with CI/CD, Git, DevOps practices, release management, and environment migration.
- Experience with Delta Live Tables, Structured Streaming, Auto Loader, Unity Catalog, and Databricks SQL.
- Exposure to data governance, lineage, cataloging, masking, encryption, and role-based access control.
- Experience in migration from legacy ETL tools or on-premise data warehouses to Databricks lakehouse platform.
- Knowledge of cloud storage services such as ADLS, Blob Storage, S3, or Google Cloud Storage.
- Experience in BFSI, healthcare, retail, or large enterprise data platform environments is an added advantage.
- Databricks, Azure, AWS, or data engineering certifications are preferred.
Pay
The expected compensation for this role ranges from $60,000 to $135,000. Final compensation will depend on various factors, including geographical location, minimum wage obligations, skills, and relevant experience.
Benefits
- Full range of medical and dental benefits options.
- Disability insurance.
- Paid time off (inclusive of sick leave).
- Other paid and unpaid leave options.