Senior Data Engineer
Tata Consultancy Services · Irvine, CA · 1 wk ago
Information Technology$120k–$180k/yrFull-time
Key Responsibilities
- Data Engineering and Development
- Design, build, test, deploy, and maintain scalable ETL and ELT pipelines using Databricks, PySpark, Spark SQL, Python, and SQL.
- Develop reusable ingestion and transformation frameworks for structured, semi-structured, and streaming data.
- Implement batch, incremental, change-data-capture, and streaming processing patterns.
- Build and maintain Delta Lake tables using medallion architecture across Bronze, Silver, and Gold layers.
- Develop dbt models, tests, macros, packages, documentation, and incremental processing patterns.
- Create and maintain Apache Airflow DAGs and Databricks Workflows with dependency management, retries, alerting, and operational controls.
- Integrate data from APIs, databases, files, event streams, and cloud data services.
- Produce technical designs, mapping specifications, lineage documentation, deployment instructions, and operational runbooks.
- Performance, Reliability, and Data Quality
- Tune Spark workloads, joins, partitioning, file sizes, caching, cluster configurations, and query plans.
- Apply Delta Lake optimization techniques, including compaction, data skipping, clustering, retention, and vacuum controls.
- Implement automated data quality, reconciliation, schema validation, observability, and freshness checks.
- Monitor pipeline health and resolve failures, performance degradation, data defects, and service-level breaches.
- Perform root-cause analysis and implement durable preventive measures.
- Support release readiness, production cutover, incident resolution, and ongoing platform operations.
- Improve compute utilization and cost efficiency across batch and streaming workloads.
- Governance, Security, and Delivery Practices
- Apply Unity Catalog standards for catalogs, schemas, tables, views, lineage, classification, and controlled access.
- Implement secure handling of credentials, secrets, personally identifiable information, and regulated data.
- Contribute to CI/CD pipelines, automated testing, code-quality checks, and environment promotion.
- Use Git-based development, peer reviews, branching standards, and release-management practices.
- Collaborate with platform engineers to deploy data assets through Terraform and Databricks Asset Bundles where applicable.
- Follow enterprise architecture, security, data-governance, and regulatory requirements.
- Collaboration and Mentoring
- Partner with architects, product owners, analysts, data scientists, governance teams, and business stakeholders.
- Translate business requirements into scalable data models, pipelines, and technical work packages.
- Conduct code reviews and enforce engineering, documentation, testing, and support standards.
- Mentor junior and mid-level engineers and share reusable patterns and best practices.
- Communicate delivery status, risks, dependencies, and technical trade-offs clearly.
Required Qualifications
- Typically 7–10 years of data engineering, data warehousing, or distributed data-processing experience.
- Strong hands-on experience with Databricks, Apache Spark, PySpark, Delta Lake, Python, and advanced SQL.
- Experience building production-grade ETL and ELT pipelines for large datasets.
- Experience with dbt Core or dbt Cloud, including models, macros, tests, documentation, and incremental processing.
- Experience with Apache Airflow, Astronomer, Databricks Workflows, or comparable orchestration platforms.
- Experience with Unity Catalog, data lineage, role-based access, and data-governance controls.
- Experience with cloud data services on AWS, Azure, or Google Cloud.
- Working knowledge of Git, CI/CD, automated testing, monitoring, and production-support practices.
- Strong troubleshooting, communication, collaboration, and technical-documentation skills.
Preferred Qualifications
- Experience in banking, financial services, insurance, asset management, risk, compliance, or another regulated industry.
- Experience modernizing Hadoop, legacy data warehouses, or traditional ETL platforms.
- Experience with Kafka, Structured Streaming, Auto Loader, Delta Live Tables, or Lakeflow Declarative Pipelines.
- Familiarity with Terraform, Databricks Asset Bundles, cloud networking, IAM, secrets management, and infrastructure automation.
- Databricks Data Engineer Associate or Professional certification.
- Experience delivering data reconciliation, regulatory reporting, test automation, and audit-ready controls.
Salary
$120,000–$180,000 a year