Senior Data Engineer
Regard · New York, NY · 3 days ago
HybridInformation Technology$165k–$220k/yrFull-time
About the role
As a Senior Data Engineer at Regard, you will own the design, development, and production deployment of the data services that power the Regard platform. From ingesting and standardizing clinical data across health systems to making it reliably available for downstream product, analytics, and machine learning workflows, you'll build and evolve the infrastructure that enables the platform.
Responsibilities
- Collect, model, and consolidate data into the data platform to support analytics, ML development, and research initiatives
- Design, build, and evolve data models and pipelines that reliably transform and deliver data to downstream consumers
- Own data quality in collaboration with engineering teams, ensuring datasets are trustworthy and production-ready
- Partner closely with product to deliver analytics and actionable insights to internal and external stakeholders
- Own the reliability and day-to-day operation of the data platform and its pipelines through proactive monitoring, alerting, and operational management
Requirements
- Bachelors degree in Computer Science, Mathematics, Statistics, or a related field, or equivalent practical experience
- 5+ years of experience in data engineering roles
- 3+ years of experience using PySpark to build data pipelines
- 3+ years of experience in public cloud provider technologies (AWS tooling such as S3, EMR, or Athena)
- Strong proficiency in Python and SQL
- Hands-on experience across the full data stack, with particular depth in data modeling and pipeline design
- Practical experience with LLM-assisted development, with an understanding of its capabilities and limitations
- Willingness to participate in on-call operational support for owned systems
Preferred Qualifications
- Experience with one or more of the following technologies: Apache Iceberg, Dagster, Clickhouse, PostgreSQL, FastAPI, Metabase
- Experience working with healthcare data, including HIPAA compliance, data de-identification, and familiarity with open data standards such as OMOP CDM
- Experience building and supporting data pipelines for ML workflows, including model training, validation, deployment, and ongoing performance evaluation