> Markdown version of [/jobs/ext/1355472-data-engineer](https://www.wearedevelopers.com/jobs/ext/1355472-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # data engineer - **Company:** Mercury - **Location:** United States (Remote available) - **Experience:** Expert - **Salary:** $105,000.0 - $115,000.0 - **Contract:** Permanent contract - **Skills:** Query Performance, Microsoft Azure, Big Data, Computer Programming, Databases, Information Engineering, Data Infrastructure, Data Structures, Electronic Data Interchange (EDI), Python (Programming Language), Machine Learning, Online Analytical Processing, Performance Tuning, Cloud Services, Tensorflow, Standard Sql, Azure Machine Learning, Data Streaming, Feature Engineering, Data Ingestion, Azure Data Factory, Pytorch, Fast Healthcare Interoperability Resources, System Availability, Kubernetes, Druid, Low Latency, Health Level Seven International, Bicep, Apache Kafka, Machine Learning Operations, Vertica, Terraform, Stream Processing, Azure Synapse Analytics, Data Pipelines, Databricks - **Published:** July 20, 2026 - **Apply:** https://www.workingnomads.com/job/go/1739413/ ## About the Role * Cloud Expertise: 5+ years of experience in data engineering, with deep proficiency in Azure Data Factory, Azure Databricks, or Azure Synapse. * OLAP Mastery: Proven experience managing and tuning ClickHouse (or similar columnar databases like Druid/Pinot) for massive datasets. * Programming: Expert-level Python and SQL skills. * ML Engineering: Familiarity with ML frameworks (PyTorch, TensorFlow) and MLOps tools (MLflow, Kubeflow, or Azure Machine Learning). * Healthcare Domain: Prior experience with healthcare data formats (HL7, FHIR, 835/837) and a strong understanding of HITRUST/HIPAA security requirements. * Scale-up Mindset: Ability to build 'v1' processes while designing for 10x growth. Preferred Qualifications: * Experience with Infrastructure as Code (Terraform, Bicep). * Knowledge of stream processing (Kafka, Azure Event Hubs). * Background in financial or payment integrity analytics. ## Description Position Summary: MedReview Innovation and Development team is seeking a data engineer to function as the primary architect and operator of our data infrastructure. Your mission is to evolve our current environment into a rapid-acquisition engine capable of feeding real-time ML models, innovation, and operations while maintaining rigorous healthcare compliance standards. Responsibilities: * Pipeline Architecture: Design, implement, and maintain end-to-end data pipelines on Azure, ensuring high availability and low latency for healthcare claim and analytics processing. * High-Performance Storage: Manage and optimize ClickHouse as our primary analytical engine, focusing on rapid data ingestion and lightning-fast query performance for large-scale datasets. * ML Data Readiness: Structure data environments to support the full ML lifecycle, from feature engineering and training to real-time model inference. * MLOps Integration: Collaborate with Data Scientists to implement automated CI/CD pipelines for model deployment, monitoring, and retraining. * Rapid Acquisition: Develop scalable frameworks to ingest diverse healthcare data sources (EDI, claims, clinical notes) with high velocity. * Security & Compliance: Ensure all data structures and processes adhere to HITRUST/HIPAA standards, collaborating with IT and the leads for technical efforts for HITRUST certification readiness. ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Back(end) to the Future: Embracing the continuous Evolution of Infrastructure and Code](https://www.wearedevelopers.com/videos/440-back-end-to-the-future-embracing-the-continuous-evolution-of-infrastructure-and-code) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Geometric deep learning for drug discovery](https://www.wearedevelopers.com/videos/264-geometric-deep-learning-for-drug-discovery) - [Modern Data Architectures need Software Engineering](https://www.wearedevelopers.com/videos/1030-modern-data-architectures-need-software-engineering) - [Serverless deployment of (large) NLP models ](https://www.wearedevelopers.com/videos/158-serverless-deployment-of-large-nlp-models) ## Related Articles - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production)