ETL Developer
Tata Consultancy Services · Plano, TX · 2 wk ago
Business Development$80k–$140k/yrFull-time
About the role
Seeking a Senior Big Data Engineer with 10–13 years of experience specializing in Hadoop, PySpark, Kafka, Hive, and strong experience designing data solutions for large-scale financial systems. The role focuses on delivering highly performant, well-governed data platforms that support the bank’s mission-critical global markets functions, including regulatory, risk, trading, and analytics workloads.
Responsibilities
- Design, develop, and optimize PySpark-based ETL pipelines running on on-prem Hadoop clusters and cloud environments.
- Build high-volume ingestion frameworks using Kafka for real-time and near-real-time trading and market data.
- Develop, tune, and manage Hadoop ecosystem components—HDFS, YARN, MapReduce, Tez, Oozie/Airflow.
- Build high-performance, optimized Hive data models for regulatory reporting, trade lifecycle, and market risk processing.
- Architect and implement Bronze/Silver/Gold layer modeling patterns within the Databricks Lakehouse.
- Apply Delta Lake best practices including:
- Optimized file management
- Z-Ordering
- Delta Change Data Feed (CDF)
- Schema evolution & enforcement
- ACID transaction handling
- Build reusable frameworks for ingestion, cleansing, transformation, and consumption of data across Lakehouse layers.
- Enable governance, lineage, and auditability using Unity Catalog or equivalent cataloging tools.
- Collaborate closely with quants, product owners, architects, risk tech, and business users.
- Participate in agile ceremonies — sprint planning, refinement, design reviews.
- Mentor junior engineers and contribute to building strong engineering practices across tech teams.
Requirements
- 10–13 years of hands-on experience in Big Data engineering.
- Expert skills in:
- PySpark — dataframe optimizations, partitioning, broadcast strategies, distributed computing.
- Kafka — producer/consumer design, schema registry, streaming ETLs.
- Hadoop ecosystem — HDFS, YARN, MapReduce/Tez, Oozie/Airflow.
- Hive — advanced query tuning, TEZ optimization, partition/bucket management.
- Extensive hands-on experience with Databricks Lakehouse, including:
- Bronze/Silver/Gold layer modeling
- Delta Lake optimizations
- Data quality frameworks on Lakehouse
- Structured & unstructured data handling
- Experience in Global Markets, Risk, Treasury, Trade Surveillance, or Regulatory Reporting.
- Strong SQL knowledge with experience working on massive datasets (TB/PB scale).
- Experience with CI/CD practices — Git, Jenkins, Bitbucket, build pipelines.
Qualifications
Bachelor of Computer Science or equivalent.
Skills
- PySpark
- Apache Kafka
- Hadoop Ecosystem
- Hive
- Databricks Lakehouse Architecture
- Delta Lake
- Bronze/Silver/Gold Data Modeling
- Big Data ETL Pipeline Development
- SQL
- Real-time Data Ingestion Frameworks
- Data Governance & Cataloging
- CI/CD Tools – Git, Jenkins, Bitbucket
- Workflow Orchestration
- Cloud & On-Prem Big Data Platforms
Benefits
- Discretionary Annual Incentive
- Comprehensive Medical Coverage:
- Medical & Health
- Dental & Vision
- Disability Planning & Insurance
- Pet Insurance Plans
- Family Support:
- Maternal & Parental Leaves
- Insurance Options:
- Auto & Home Insurance
- Identity Theft Protection
- Convenience & Professional Growth:
- Commuter Benefits
- Certification & Training Reimbursement
- Time Off:
- Vacation
- Time Off
- Sick Leave & Holidays
- Legal & Financial Assistance:
- Legal Assistance
- 401K Plan
- Performance Bonus
- College Fund
- Student Loan Refinancing
Pay
$80,000–140,000 a year