> Markdown version of [/jobs/ext/1226474-sr-data-engineer](https://www.wearedevelopers.com/jobs/ext/1226474-sr-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Sr Data Engineer - **Company:** Honeywell International Inc. - **Location:** Atlanta, GA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Training Data, Agile Methodology, Artificial Intelligence, Microsoft Azure, Databases, Continuous Integration, Data Architecture, Information Engineering, Data Security, Data Systems, Software Debugging, Distributed Computing Environment, Distributed Systems, Github, Machine Learning, Operational Databases, Software Product Management, DataOps, Data Streaming, Data Processing, Google Cloud, Feature Engineering, Retrieval-Augmented Generation, Large Language Models, Apache Spark, Generative AI, Containerization, Data Lakes, Pyspark, Git Flow, Kubernetes, Data Lineage, Data Analytics, Performance Monitor, Apache Kafka, Data Management, Machine Learning Operations, Stream Processing, Stream Analytics, Data Pipelines, Docker, Databricks - **Published:** July 10, 2026 - **Apply:** https://www.juju.com/job/00000000geyelv ## About the Role + Minimum 5 years of experience building production data pipelines in Databricks processing TB scale data + Extensive experience implementing medallion architecture (Bronze/Silver/Gold) with Delta Lake, Delta Live Tables (DLT), and Lakeflow for batch and streaming pipelines from + Event Hub or Kafka sources + Strong hands-on proficiency with PySpark for distributed data processing and transformation + Strong experience working with cloud platforms such as Azure, GCP and Databricks, especially in designing and implementing AI/ML-driven data workflows + Proficient in CI/CD practices using Databricks Asset Bundles (DAB), Git workflows, GitHub Actions, and understanding of DataOps practices including data quality testing and observability + Hands-on experience building RAG applications with vector databases, LLM integration, and agentic frameworks like LangChain, LangGraph + Natural analytical mindset with demonstrated ability to explore data, debug complex distributed systems, and optimize pipeline performance at scale WE VALUE + Experience building RAG and agentic architecture solutions and working with LLM-powered applications + Expertise in real-time data processing frameworks (Apache Spark Streaming, Structured Streaming) + Knowledge of MLOps practices and experience building data pipelines for AI model deployment + Experience with time-series databases and IoT data modeling patterns + Familiarity with containerization (Docker) and orchestration (Kubernetes) for AI workloads + Strong background in data quality implementation for AI training data + Experience working with distributed teams and cross-functional collaboration + Knowledge of data security and governance practices for AI systems + Experience working on analytics projects with Agile and Scrum Methodologies US PERSON REQUIREMENT Due to compliance with U.S. export control laws and regulations, candidate must be a U.S. Person, which is defined as a U.S. citizen, a U.S. permanent resident, or have protected status in the U.S. under asylum or refugee status, or have the ability to obtain an export authorization. ## Description As a Senior Data Engineer, you will be part of a high-performing global team delivering advanced AI and data solutions for Honeywell's industrial customers, with a focus on IoT and real-time data processing. In this role, you will design and implement scalable data architectures and pipelines that enable next-generation AI capabilities, including large-scale machine learning models, intelligent automation, and real-time analytics. You will work closely with cross-functional teams to transform high-volume IoT telemetry into reliable, actionable insights that support Honeywell's connected industrial solutions., You will report directly to our Data Engineering Manager and you'll work out of our Atlanta, GA location on a Hybrid work schedule. Note: for the first 90 days, new hires must be prepared to work 100% onsite M-F., Data Engineering & AI Pipeline Development: + Design and implement scalable data architectures to process high-volume IoT sensor data and telemetry streams, ensuring reliable data capture and processing for AI/ML workloads + Build and maintain data pipelines for AI product lifecycle, including training data preparation, feature engineering, and inference data flows + Develop and optimize RAG (Retrieval Augmented Generation) systems, including vector databases, embedding pipelines, and efficient retrieval mechanisms + Lead the architecture and development of scalable data platforms on Databricks + Drive the integration of GenAI capabilities into data workflows and applications + Optimize data processing for performance, cost, and reliability at scale + Create robust data integration solutions that combine industrial IoT data streams with enterprise data sources for AI model training and inference DataOps: + Implement DataOps practices to ensure continuous integration and delivery of data pipelines powering AI solutions + Design and maintain automated testing frameworks for data quality, data drift detection, and AI model performance monitoring + Create self-service data assets enabling data scientists and ML engineers to access and utilize data efficiently + Design and maintain automated documentation for data lineage and AI model provenance Collaboration & Innovation: + Partner with ML engineers and data scientists to implement efficient data workflows for model training, fine-tuning, and deployment + Mentor team members and provide technical leadership on complex data engineering challenges + Establish data engineering best practices, including modular code design and reusable frameworks + Drive projects to completion while working in an agile environment with evolving requirements in the rapidly changing AI landscape ## Related Videos - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Empowering Retail Through Applied Machine Learning](https://www.wearedevelopers.com/videos/976-empowering-retail-through-applied-machine-learning) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) ## Related Articles - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [What Industries Outside of AI Are Hiring The Most AI Experts?](https://www.wearedevelopers.com/magazine/98-what-industries-outside-of-ai-are-hiring-the-most-ai-experts)