Staff Software Engineer, Data Products
About the role
Omada Health is on a mission to bend the curve of chronic disease. We rely on trusted, high-quality data to power intelligent products, personalized member experiences, and data-driven decision making. As machine learning becomes increasingly central to our platform, we're investing in the data foundations that enable scalable model development, experimentation, and production inference.
We are seeking a Staff Software Engineer, Data Engineering to lead the design and development of the production data platform that powers machine learning across Omada. In this role, you will partner closely with Data Scientists, Applied AI Engineers, Product Engineers, and fellow Data Engineers to identify, design and build trusted, reusable datasets foundations that serve as the foundation for feature engineering, model training, experimentation, and production inference.
Responsibilities
- Design, build, and maintain reusable feature datasets that support machine learning use cases including personalization, engagement, risk prediction, churn modeling, recommendation systems, and experimentation.
- Establish self-service foundations that streamline and democratize dataset creation across the data organization.
- Partner with Data Scientists to translate modeling requirements into production-ready feature pipelines, supporting the full model lifecycle from exploration to deployment.
- Identify source data, transformations, and historical windows needed for feature engineering.
- Help define and build shared, reusable feature definitions across models rather than one-off datasets.
- Balance features freshness, correctness, latency, and computational efficiency when designing data pipelines.
- Build datasets that support both historical model training and future production inference.
- Design and implement batch and streaming pipelines that transform raw healthcare, behavioral, product, and operational data into trusted ML-ready datasets.
- Build reliable data processing systems using Python, SQL, Spark, and modern cloud data platforms.
- Optimize large-scale distributed processing for performance, scalability, and cost.
- Design data pipelines that are modular, testable, observable, and easy to evolve as product requirements change.
- Ensure data quality through testing, anomaly detection, schema validation, and pipeline monitoring.
- Partner with platform teams to support near real-time feature generation where appropriate.
- Improve reproducibility by standardizing feature computation across experimentation and production.
- Support rapid experimentation without sacrificing long-term maintainability.
- Ensure data quality through testing, anomaly detection, schema validation, and pipeline monitoring.
- Establish engineering standards for correctness, documentation, and maintainability.
Requirements
- 8+ years building large-scale production data platforms and distributed data pipelines.
- Experience designing reusable datasets that power machine learning, experimentation, or advanced analytics.
- Demonstrated experience partnering closely with Data Scientists to productionize feature engineering workflows.
- Experience leading cross-team technical initiatives and influencing engineering direction.
- Strong experience working with cloud-native data platforms such as AWS.
- Experience building production data systems using Databricks, Iceberg, Spark, Redshift, Snowflake, or similar technologies.
- Experience developing reliable batch and streaming data pipelines.
- Experience working with healthcare, behavioral, or other large-scale event data is a plus.
Technical Skills
- Expert SQL with strong data modeling skills.
- Strong programming skills in Python, Java, or Scala.
- Experience with Apache Spark or similar distributed compute frameworks.
- Experience with Airflow or similar orchestration platforms.
- Experience designing dimensional models, event models, and feature datasets.
- Experience implementing testing, CI/CD, observability, and production monitoring for data pipelines.
- Understanding of Feature Stores and ML data lifecycle concepts.
- Experience with Lakehouse Architecture such as Databricks, Iceberg is a strong plus.
- Experience with streaming technologies such as Kafka, Flink, or Spark Structured Streaming.
- Familiarity with NoSQL Databases (document & graph databases Nepture, Neo4j etc.)
- Understanding of software engineering best practices, distributed systems, and cloud-native architectures.
Communication Skills
An exceptional people leader who develops engineers into future technical leaders. Comfortable influencing senior executives and cross-functional partners. Skilled at balancing business priorities with long-term technical investments. Able to communicate complex technical concepts to both technical and non-technical audiences. Passionate about building trusted data platforms that enable the business.
Education
Bachelor’s degree in Computer Science or a similar discipline preferred.
Technologies we use
Ruby on Rails, Redshift, Athena, Postgres, SQL, Python, Apache Airflow, Appflow, S3, SNS, SQS, Kafka, Docker, Kubernetes, AWS infrastructure, Lambda, Serverless, Tableau, Bugsnag, Datadog, GitLabCI, Cursor, OpenMetadata, Databricks
Benefits
- Competitive salary with generous annual cash bonus
- Equity grants
- Remote first work from home culture
- Flexible Time Off to help you rest, recharge, and connect with loved ones
- Generous parental leave
- Health, dental, and vision insurance (and above market employer contributions)
- 401k retirement savings plan
- Lifestyle Spending Account (LSA)
- Mental Health Support Solutions
Pay
These ranges represent a good faith estimate of the minimum and maximum base salary for the position in that zone at the time of posting. Zone 1: $202,400 - $253,000 Zone 2: $193,600 - $242,000 Zone 3: $176,000 - $220,000