Jobs · Information Technology

Data Engineer - Remote

UnitedHealthcare · Minnetonka, MN · 2 wk ago
Information Technology$73k–$130k/yrFull-time

At UnitedHealthcare, we're simplifying the health care experience, creating healthier communities and removing barriers to quality care. The work you do here impacts the lives of millions of people for the better. Come build the health care system of tomorrow, making it more responsive, affordable and optimized.

About the role

This role is responsible for designing, developing, and maintaining scalable and reliable data pipelines that support both batch and real-time analytics within an Azure-based data platform. The position operates as part of a collaborative data engineering team, working closely with fellow data engineers and data science & reporting partners to meet evolving data requirements.

You'll enjoy the flexibility to work remotely from anywhere within the U.S. For hires in the Minneapolis or Washington, D.C. area, you will be required to work in the office a minimum of four days per week.

Responsibilities

  • Design, develop, and maintain robust pipelines to ingest data from various sources (both streaming and batch) into the analytics environment using Azure Data Factory and PySpark via Databricks.
  • Set up real-time data ingestion using tools like Spark Structured Streaming and batch ETL jobs for periodic data loads. Ensure these pipelines are scalable, efficient, and fault-tolerant to handle growing data volumes and velocity.
  • Implement Data as per Medallion Architecture: Utilize the Medallion (Bronze/Silver/Gold) architecture principles to organize data processing stages. Establish raw data capture (bronze), perform cleansing and transformations (silver), and curate refined datasets for analysis and machine learning (gold).
  • Apply best practices in each layer, such as schema enforcement and checkpointing for streaming data.
  • Optimize Spark jobs by tuning configurations, improving query logic, and managing resources to achieve high throughput and low latency. Address bottlenecks in streaming pipelines (e.g., by scaling clusters or tweaking batch intervals) and ensure timely data delivery.
  • Optimize job scheduling and cluster utilization to balance timely data delivery with cost-effectiveness.
  • Build and maintain data pipelines with an emphasis on data cleaning steps. Integrate data from various sources (APIs, databases, file feeds, IoT streams, etc.) into the data platform, writing transformations that handle anomalies (e.g., missing or corrupt values) and standardize datasets.
  • Collaborate with other data engineers to share responsibility across different pipelines or sources, ensuring redundancy and knowledge transfer.
  • Implement comprehensive data validation rules and checks within pipelines. For example, verify schema correctness, check value ranges for sensor or health data, and ensure referential integrity where applicable.
  • Set up automated alerts or logs that flag inconsistent or bad data, enabling quick intervention. Build a library of data quality tests that run as part of the pipeline (for both streaming and batch processes) to catch issues early.
  • Leverage modern pipeline frameworks and tools to improve development productivity. For example, use Databricks Delta Live Tables or Lakehouse pipelines to declaratively define data flows where applicable.
  • Explore the use of Spark Declarative Lakeflow Pipelines or similar technologies to simplify the orchestration of complex data processes.
  • Implement monitoring and alerting for pipeline health. Investigate and resolve problems such as data delays, pipeline failures, or data inconsistencies.
  • Use logs, error messages, and analytics to identify root causes (e.g., source system changes, bug in transformation logic) and implement fixes.
  • Work closely with other data engineers and data science team members to understand data requirements and adjust pipelines accordingly.
  • Document data engineering workflows and ensure proper data governance (security, privacy, access controls) is in place.
  • Maintain clear documentation of data pipelines, including data source details, transformation logic, and data destination schemas.
  • Ensure that data lineage is tracked so one can trace how data moved and changed through the system.
  • Adhere to data governance policies—for instance, ensure sensitive data is properly masked or encrypted in non-production environments, and that access controls are in place.
  • Work with leadership to periodically review and improve data management practices.

Requirements

  • 3+ years of experience in a data engineering role designing and implementing data pipelines and ETL processes. Understanding of how to handle incremental data loads and maintain history (CDC - change data capture).
  • 2+ years of experience in SQL for data manipulation and query optimization.
  • Knowledge of Python and Apache Spark (using PySpark) for building data pipelines; ability to write efficient code for batch and streaming data transformations.
  • 1+ years of experience using Azure, Databricks, or an equivalent cloud-based data platform. Comfortable with managing clusters, using notebooks, and working with Delta Lake or Parquet files.
  • Familiarity with cloud data services and tools for pipeline orchestration is expected.
  • Experience working in a team environment with agile methodologies.
  • Ability to communicate effectively with both technical peers and non-technical stakeholders (explaining data issues in plain language).
  • Comfortable using version control systems and participating in collaborative development (code reviews, pair programming when needed).
  • Familiar with streaming data technologies such as Spark Streaming, Kafka, Azure Event Hubs, or similar platforms for real-time data ingestion.
  • Demonstrated ability to detect and correct data issues—for instance, identifying when a data source has stopped updating, or when an upstream change has altered data format.
  • Experience implementing validation checks or using frameworks to enforce data quality standards.

Qualifications

  • Experience with any declarative pipeline frameworks or data workflow management tools (e.g., Databricks Delta Live Tables).
  • Experience integrating data quality checks into pipelines (such as using assertions or Great Expectations tests) to ensure accuracy and completeness of data.
  • Familiarity with data security practices, encryption, and handling of sensitive data.
  • Demonstrated skill in performance tuning for Spark or SQL queries. For example, experience in partitioning strategies, caching, or troubleshooting shuffle issues to optimize heavy data workloads.

Benefits

In addition to your salary, we offer benefits such as a comprehensive benefits package, incentive and recognition programs, equity stock purchase, and 401k contribution (all benefits are subject to eligibility requirements).

Pay

The salary for this role will range from $72,800 to $130,000 annually based on full-time employment. Pay is based on several factors including but not limited to local labor markets, education, work experience, and certifications.

Similar jobs

Data Engineer - Remote

SundayyUnited States· 1 wk ago
RemoteInformation Technology$99k–$203k/yrapply on sundayy.com

Data Engineer - Remote

Sentara HealthVirginia Beach, VA· 1 mo ago
Information Technology$80k–$134k/yrapply on sentara.wd1.myworkdayjobs.com

Data Engineer | Remote

CodeGeniusRecruitUnited States· 6 days ago
RemoteInformation Technology$80/hrapply on candidateportal.ceipal.com

Data Engineer | Remote

Crossing HurdlesUnited States· 1 mo ago
RemoteInformation Technology$140k–$180k/yrapply on candidateportal.ceipal.com

Data Engineer – Remote

YO IT ConsultingLos Angeles, CA· 3 mo ago
RemoteInformation Technologyapply on yohrconsultancy.hiresome.ai