Jobs · Information Technology · Texas

Data Engineer

Tata Consultancy Services · Irving, TX · 2 days ago
Information Technology$125k–$140k/yrFull-time

About the role

We are seeking a highly skilled and motivated Data Engineer to play a pivotal role in designing, building, and optimizing our next-generation scalable data pipelines. This position requires expertise in processing massive datasets using cutting-edge technologies like Apache Spark, PySpark, and Hive within Cloudera Platform. Your primary objective will be to ensure the utmost data reliability, speed, and efficiency, providing a robust foundation for downstream business intelligence and advanced analytics initiatives.

Responsibilities

  • Design, build, and maintain highly scalable and efficient ETL/ELT data pipelines utilizing PySpark and Spark SQL, Hive for complex data transformations.
  • Strategically manage data layout, partitioning, and indexing within Apache Hive and various cloud data lake solutions to optimize performance and accessibility.
  • Proactively identify and resolve performance bottlenecks in Spark jobs, leveraging Spark UI for in-depth analysis, effectively managing data skewness, and optimizing memory utilization.
  • Develop robust solutions for ingesting high-volume and diverse datasets from both structured relational databases and unstructured flat files into our data ecosystem.
  • Implement and manage automated data workflows using industry-standard scheduling tools like Apache Airflow or platform-native schedulers, ensuring timely and reliable data delivery.
  • Partner closely with data scientists, business analysts, and cross-functional enterprise teams to translate complex business requirements into technically sound and efficient data solutions.

Requirements

  • Big Data Frameworks Expertise: Demonstrated high proficiency in Apache Spark architecture, including a deep understanding of drivers, executors, and Directed Acyclic Graphs (DAGs).
  • Advanced Programming: Exceptional coding skills in Python and extensive experience with the PySpark API for developing intricate data transformations and processing logic.
  • Querying & Schema Management: Strong command of HiveQL and ANSI SQL, coupled with expertise in data partitioning techniques and effective schema definition.
  • Optimized Storage Formats: In-depth understanding and practical experience with optimized big data storage file formats such as Parquet, ORC, and Avro.
  • Data Warehousing Fundamentals: Solid foundation in Dimensional Data Modeling, including Star and Snowflake schemas, and practical experience with Data Lakes concepts and implementation.
  • Bachelor of Computer Science.

Preferred Qualifications

  • CI/CD & DevOps Automation: Experience with Continuous Integration/Continuous Deployment (CI/CD) practices and automation tools like Git, Jenkins, or Ansible.
  • Cloud Ecosystem Development: Experience in development utilizing cloud-native big data utilities (e.g., AWS EMR, AWS Databricks) within major cloud platforms.
  • NoSQL Database Integration: Exposure to and experience with NoSQL databases such as HBase, Cassandra, or MongoDB.
  • Professional Certifications: Relevant professional certifications on Spark or Data Engineer are highly valued.

Pay

$125,000 to $140,000 per year

Similar jobs

Data Engineer

IBMYorktown Heights, NY· 2 wk ago
Information Technologyapply on ibmglobal.avature.net

Data Engineer

DLA PiperHouston, TX· 2 wk ago
Information Technology$101k–$160k/yrapply on dlapiper.wd1.myworkdayjobs.com

Data Engineer

CodeVertex TalentOhio, United States· 2 wk ago
RemoteAnalyst$125k–$180k/yrapply on tally.so

Data Engineer

CodeVertex SystemNorth Carolina, United States· 2 wk ago
RemoteAnalyst$125k–$180k/yrapply on tally.so

Data Engineer

HaystackCalifornia, United States· 2 wk ago
Information Technology$99k–$225k/yrapply on haystack.cv

Data Engineer

Security BenefitTopeka, KS· 2 wk ago
$60/hrapply on jobs.dayforcehcm.com