> Markdown version of [/jobs/ext/1853605-remote](https://www.wearedevelopers.com/jobs/ext/1853605-remote). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Remote - **Company:** Omada Health, Inc. - **Location:** United States (Remote available) - **Experience:** Expert - **Salary:** $176,000.0 - $253,000.0 - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Airflow, Amazon Web Services, Amazon S3, Business Logic, Cloud Engineering, Encodings, Computer Programming, Continuous Integration, Data Architecture, Information Engineering, Data Systems, Cursor (Graphical User Interface Elements), Distributed Computing Environment, Distributed Data Store, Distributed Systems, Graph Database, Python (Programming Language), PostgreSQL, Machine Learning, Neo4j, NoSQL, Operational Data Store, Operational Databases, Ruby on Rails, Recommender Systems, Cloud Services, Amazon Simple Notification Service (SNS), Software Construction, SQL Databases, Data Streaming, Tableau (Software), Datadog, Data Processing, Feature Engineering, Snowflake, Apache Spark, Gitlab-ci, Kubernetes, Information Technology, Apache Flink, Apache Kafka, Spark Streaming, Data Management, Machine Learning Operations, Video Streaming, Functional Programming, Amazon Simple Queue Service (SQS), Data Pipelines, Serverless Computing, Docker, Amazon Redshift, Databricks - **Published:** July 21, 2026 - **Apply:** https://job-boards.greenhouse.io/omadahealth/jobs/8073030 ## About the Role * 8+ years building large-scale production data platforms and distributed data pipelines. * Experience designing reusable datasets that power machine learning, experimentation, or advanced analytics. * Demonstrated experience partnering closely with Data Scientists to productionize feature engineering workflows. * Experience leading cross-team technical initiatives and influencing engineering direction. * Strong experience working with cloud-native data platforms such as AWS. * Experience building production data systems using Databricks, Iceberg, Spark, Redshift, Snowflake, or similar technologies. * Experience developing reliable batch and streaming data pipelines. * Experience working with healthcare, behavioral, or other large-scale event data is a plus. Technical Skills * Expert SQL with strong data modeling skills. * Strong programming skills in Python, Java, or Scala. * Experience with Apache Spark or similar distributed compute frameworks. * Experience with Airflow or similar orchestration platforms. * Experience designing dimensional models, event models, and feature datasets. * Experience implementing testing, CI/CD, observability, and production monitoring for data pipelines. * Understanding of Feature Stores and ML data lifecycle concepts. * Experience with Lakehouse Architecture such as Databricks, Iceberg is a strong plus. * Experience with streaming technologies such as Kafka, Flink, or Spark Structured Streaming. * Familiarity with NoSQL Databases (document & graph databases Nepture, Neo4j etc.) * Understanding of software engineering best practices, distributed systems, and cloud-native architectures. Communication Skills: An exceptional people leader who develops engineers into future technical leaders. * Comfortable influencing senior executives and cross-functional partners. * Skilled at balancing business priorities with long-term technical investments. * Able to communicate complex technical concepts to both technical and non-technical audiences. * Passionate about building trusted data platforms that enable the business. Education: Bachelor's degree in Computer Science or a similar discipline preferred. Technologies we use: Ruby on Rails, Redshift, Athena, Postgres, SQL, Python, Apache Airflow, Appflow, S3, SNS, SQS, Kafka, Docker, Kubernetes, AWS infrastructure, Lambda, Serverless, Tableau, Bugsnag, Datadog, GitLabCI, Cursor, OpenMetadata, Databricks, * Experience supporting personalization, recommendation, ranking, or predictive modeling systems. * Familiarity with model training pipelines and MLOps workflows. * Experience designing data platforms for experimentation. * Healthcare industry experience is a plus. * Experience with Data/AI Governance. ## Description We are seeking a Staff Software Engineer, Data Engineering to lead the design and development of the production data platform that powers machine learning across Omada. In this role, you will partner closely with Data Scientists, Applied AI Engineers, Product Engineers, and fellow Data Engineers to identify, design and build trusted, reusable datasets foundations that serve as the foundation for feature engineering, model training, experimentation, and production inference. Rather than building one-off pipelines for individual models, you'll create scalable data products and feature pipelines that enable multiple machine learning use cases while ensuring consistency, reliability, and governance across the ML lifecycle. You will own the technical design of feature datasets-from ingesting raw behavioral, clinical, and operational data through transforming, validating, and publishing production-grade datasets that are reusable across modeling teams. This role is ideal for someone who enjoys solving complex data problems, designing scalable distributed data systems, and enabling machine learning through well-engineered data foundations., * Design, build, and maintain reusable feature datasets that support machine learning use cases including personalization, engagement, risk prediction, churn modeling, recommendation systems, and experimentation. * Establish self-service foundations that streamline and democratize dataset creation across the data organization. * Partner with Data Scientists to translate modeling requirements into production-ready feature pipelines, supporting the full model lifecycle from exploration to deployment. * Identify source data, transformations, and historical windows needed for feature engineering. Help define and build shared, reusable feature definitions across models rather than one-off datasets. * Balance features freshness, correctness, latency, and computational efficiency when designing data pipelines. * Build datasets that support both historical model training and future production inference. * Design and implement batch and streaming pipelines that transform raw healthcare, behavioral, product, and operational data into trusted ML-ready datasets. * Build reliable data processing systems using Python, SQL, Spark, and modern cloud data platforms. * Optimize large-scale distributed processing for performance, scalability, and cost. * Design data pipelines that are modular, testable, observable, and easy to evolve as product requirements change. * Ensure data quality through testing, anomaly detection, schema validation, and pipeline monitoring. * Partner with platform teams to support near real-time feature generation where appropriate. * Improve reproducibility by standardizing feature computation across experimentation and production. * Support rapid experimentation without sacrificing long-term maintainability. * Ensure data quality through testing, anomaly detection, schema validation, and pipeline monitoring. * Establish engineering standards for correctness, documentation, and maintainability. * Familiarity with feature stores or feature management platforms. * Familiarity with model training pipelines and MLOps workflows. Technical Leadership: * Lead architecture and design discussions for large-scale ML data systems, driving adoption of reusable patterns and platform capabilities across Data Engineering. * Influence technical direction across multiple engineering teams, embedding with Product, Engineering, and business stakeholders (Clinical, Finance, Growth, Enrollment) during early design phases to shape data capture requirements at the source. * Translate ambiguous business requirements from Business domain SMEs into concrete technical specs, maintaining consistency of business logic and definitions across systems. * Mentor engineers on distributed data processing, software engineering best practices, and scalable data modeling. ## Related Videos - [Putting the Graph In GraphQL With The Neo4j GraphQL Library](https://www.wearedevelopers.com/videos/257-putting-the-graph-in-graphql-with-the-neo4j-graphql-library) - [Leveraging Real time data in FSIs](https://www.wearedevelopers.com/videos/806-leveraging-real-time-data-in-fsis) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [Cyber Sleuth: Finding Hidden Connections in Cyber Data](https://www.wearedevelopers.com/videos/893-cyber-sleuth-finding-hidden-connections-in-cyber-data) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk)