> Markdown version of [/jobs/ext/2169143-data-engineer](https://www.wearedevelopers.com/jobs/ext/2169143-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Engineer - **Company:** Ngtalenttech Group Llc - **Location:** United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Unity 3d, Java (Programming Language), Microsoft Excel, Adobe InDesign, Application Programming Interfaces (APIs), Artificial Intelligence, Airflow, Amazon Web Services, Amazon S3, Business Analytics Applications, Cloud Engineering, Continuous Integration, Data as a Services, Data Architecture, Information Engineering, Data Infrastructure, Extract Transform Load (ETL), Data Security, Data Systems, Distributed Computing Environment, Amazon DynamoDB, Electronic Data Interchange (EDI), Python (Programming Language), Key Management, NoSQL, Software Tools, Salesforce.Com, SQL Databases, Management of Software Versions, Data Logging, Data Processing, Real Time Systems, Feature Engineering, Chatbots, Data Ingestion, Delivery Pipeline, Large Language Models, Apache Spark, State Machines, Electronic Medical Records, Indexer, Data Layers, Data Lakes, Pyspark, Kubernetes, Apache Kafka, Machine Learning Operations, Hubspot, Stream Processing, Data Pipelines, Docker, Databricks - **Published:** August 21, 2026 - **Apply:** https://www.dice.com/job-detail/482a8809-9d0c-4860-a51b-357650ef48b1 ## About the Role 7-10 years of professional data engineering experience, including 3+ years with modern cloud-based data lake or lakehouse architectures. Deep expertise in Apache Spark (PySpark, Scala, or Java) for large-scale distributed data processing and strong hands-on Databricks experience across Delta Lake, Unity Catalog, and Workflows. Hands-on experience with AWS data services: S3, Redshift, Athena, Glue, Lambda, Step Functions, Kinesis, DynamoDB, EMR, Bedrock, Transcribe, and Comprehend Medical. Experience in healthcare or regulated data environments; familiarity with HIPAA technical safeguards required. Demonstrated experience building data infrastructure in support of ML feature stores and model serving pipelines. Technical Skills Strong Python (preferred); Java or Scala also valued. Proficient SQL across relational, NoSQL, and lakehouse environments. ETL/ELT orchestration with Airflow, dbt, and AWS Step Functions; stream processing with Kafka/Kinesis. Experience with vector databases, embedding stores, document ingestion pipelines, and RAG pipeline data layers for LLM use cases. Container orchestration with Docker and Kubernetes/EKS; CI/CD for data pipelines; familiarity with data quality and data contract frameworks. BI tool integration, ad-hoc query execution engine design, query result caching, and multi-format report delivery pipelines. Soft Skills Strong communicator able to translate data architecture decisions to clinical, operational, and executive stakeholders. Collaborative by nature; comfortable driving alignment across data science, product, engineering, and clinical teams. Detail-oriented and committed to data quality, HIPAA compliance, and operational excellence; self-directed with a continuous learning mindset. ## Description The Senior Data Engineer will be a cornerstone of the data platform team, architecting and operating the cloud-native, lakehouse-based infrastructure that powers our full suite of AI and analytics products - from clinical and caregiver risk intelligence to LLM-enabled chatbots, real-time signal capture, and payer-facing outcome reporting. The ideal candidate brings deep, hands-on expertise with Databricks, Apache Spark, and AWS Cloud services, and can serve as both a technical leader and mentor. You will own end-to-end pipeline design from raw source ingestion through curated analytics-ready datasets, ensuring data systems are reliable, scalable, and HIPAA-compliant., Architect and build robust, scalable, and secure data pipelines leveraging Databricks (Delta Lake, Spark, Unity Catalog, Workflows) and AWS (S3, Redshift, Athena, Glue, Lambda, Step Functions, Kinesis), supporting both batch and real-time processing modes. Design and evolve the enterprise lakehouse architecture, including data modeling, Delta table optimization, data versioning, and a semantic data layer and metrics store serving BI tools and downstream ML feature engineering. Build and maintain ETL/ELT orchestration (Airflow, dbt) and stream processing pipelines (Kafka/Kinesis) ingesting claims, EHR/EMR, pharmacy, IoT, survey, CRM, and digital engagement event data. Partner with Data Scientists to design and maintain feature store architecture, real-time member context assembly, and temporal feature computation services supporting risk and predictive models. Develop data infrastructure for LLM and NLP use cases: vector databases, knowledge base ingestion, transcript storage and indexing, secure audio storage, and LLM response logging and audit trails. Build APIs and integrations for downstream consumers: care team dashboards, lead scoring, pathway recommendation, queue prioritization, CRM connectors (Salesforce, HubSpot), and payer-facing data exchange. Implement HIPAA-compliant data handling across all pipelines including PHI detection, redaction, audit trails, row-level security, SSO, and secrets management (AWS Secrets Manager, Vault). Build reporting infrastructure including a report template engine, multi-format export (PDF, Excel, CSV), scheduled report generation, and delivery APIs for internal and payer-facing reporting. Conduct pull request reviews, enforce engineering best practices, mentor junior engineers, and research and apply emerging data engineering tools and patterns. Participate in design discussions with technical leads across product lines; perform other duties and special projects as assigned. ## Related Videos - [Integrate your Cognitive Assistant with 3rd-party DBs and software](https://www.wearedevelopers.com/videos/249-integrate-your-cognitive-assistant-with-3rd-party-dbs-and-software) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Leveraging Real time data in FSIs](https://www.wearedevelopers.com/videos/806-leveraging-real-time-data-in-fsis) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) - [Modern Data Architectures need Software Engineering](https://www.wearedevelopers.com/videos/1030-modern-data-architectures-need-software-engineering) ## Related Articles - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [The Most Popular IT Jobs on the Market](https://www.wearedevelopers.com/magazine/376-the-most-popular-it-jobs-on-the-market)