Data Engineer
Pipeline Engineering (ETL/ELT)
Architect, write, and deploy resilient batch and streaming data pipelines using Python or Scala to extract, transform, and load data from transactional databases, webhooks, and third-party SaaS APIs.
Data Warehouse & Lakehouse Optimization
Design and optimize schemas, tables, and partitions within cloud data warehouses (e.g., Snowflake, BigQuery, Redshift) to reduce query latencies and maximize storage efficiency.
Data Quality & Observability Automation
Implement automated testing frameworks and monitoring alerts (using tools like Great Expectations or Datadog) to proactively detect data drift, schema changes, and pipeline anomalies.
Workflow Orchestration Management
Author, schedule, and maintain complex, dependency-aware Directed Acyclic Graphs (DAGs) using workflow management tools like Apache Airflow, Prefect, or Dagster.
Data Modeling & Governance
Collaborate closely with backend engineers and analytics teams to design logical data models (e.g., Star Schema, Data Vault 2.0), enforce data governance policies, and ensure strict compliance with global security standards.
Requirements
- Educational Background: Bachelor’s degree in Computer Science, Computer Engineering, Information Systems, or an equivalent technical discipline with a strong emphasis on distributed systems and databases.
- Professional Experience: 2 to 5 years of proven professional experience as a dedicated data engineer or backend platform developer operating in a production cloud environment.
- Technical Proficiency: Advanced SQL Command, expert-level mastery of SQL, with a deep understanding of query optimization, window functions, complex indexing, and profiling execution plans; Strong proficiency writing clean, testable production code in Python, Scala, or Java, applying object-oriented or functional programming principles; Practical, hands-on experience utilizing core services within major cloud ecosystems—specifically AWS or GCP (e.g., S3/GCS, EC2/Compute Engine, IAM, Cloud Functions); Foundational knowledge of distributed data processing concepts and storage architectures (e.g., Apache Spark, Hadoop, Parquet formats).
Preferred Qualifications
- Prior experience utilizing data build tools for managing analytics engineering transformations.
- Hands-on experience with streaming architectures and message brokers (such as Apache Kafka, Amazon Kinesis, or RabbitMQ).
- Familiarity with infrastructure-as-code software (Terraform) and container orchestration tools (Docker, Kubernetes).