Data Engineer
Scale.jobs · New York, NY · 1 mo ago
RemoteRemoteEngineeringFull-time
About The Role
The role owns the architecture, development, and maintenance of scalable data pipelines and storage systems. This position bridges the gap between raw backend application data and the analytical structures required for business intelligence, data science, and operational reporting. The engineer will collaborate closely with software engineers, product managers, and data analysts to design robust data models, optimize query performance, and ensure absolute data integrity and availability across the entire modern data stack.
Key Responsibilities
- Design, build, and optimize scalable ELT/ETL pipelines using Python, SQL, and orchestration tools like Apache Airflow or Prefect
- Architect and maintain data warehouse schemas in Snowflake, BigQuery, or Redshift, implementing clean data modeling practices like dbt, Star Schema, or Data Vault
- Implement data quality monitoring, anomaly detection, and automated testing frameworks to ensure high reliability of critical data assets
- Optimize query performance, storage costs, and database indexing strategies across multi-terabyte data warehouses
- Collaborate with backend engineering teams to ingest data from relational databases, NoSQL stores, and third-party APIs via Kafka, Kinesis, or batch processes
- Establish infrastructure-as-code for data platforms using Terraform and manage containerized data workloads with Docker and Kubernetes
What We Are Looking For
- 3-6 years of experience in data engineering, software engineering, or a highly quantitative analytical role
- Expert-level SQL proficiency and strong programming skills in Python or Scala for complex data manipulation
- Production experience with dbt (data build tool) and a modern cloud data warehouse such as Snowflake, BigQuery, or Databricks
- Hands-on experience with workflow orchestration platforms like Airflow, Dagster, or Prefect, and cloud infrastructure platforms (AWS, GCP, or Azure)
- Solid understanding of data warehousing concepts, dimensional modeling, and database tuning
- Bonus: Experience with streaming data technologies (Kafka, Flink), Infrastructure as Code (Terraform), or managing BI tool semantic layers (Looker, Tableau)