> Markdown version of [/jobs/ext/2999323-lead-data-engineer-physical-ai-platform-data-engineering](https://www.wearedevelopers.com/jobs/ext/2999323-lead-data-engineer-physical-ai-platform-data-engineering). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Lead Data Engineer - Physical AI Platform, Data Engineering - **Company:** Caterpillar - **Location:** Chicago, IL, United States - **Experience:** Expert - **Salary:** $128,470.0 - $208,770.0 - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Application Programming Interfaces (APIs), Amazon Web Services, Amazon S3, JIRA, Automation of Tests, Microsoft Azure, Big Data, Databases, Computer Engineering, Data as a Services, Data Architecture, Information Engineering, Data Infrastructure, Data Integration, Data Integrity, Data Systems, Distributed Data Store, Amazon DynamoDB, Electronic Data Interchange (EDI), Python (Programming Language), NoSQL, Operational Databases, Performance Tuning, Software Tools, Cloud Services, Software Deployment, Software Engineering, SQL Databases, Data Streaming, Systems Integration, Strategies of Testing, Web Application Frameworks, Data Processing, Enterprise Software Applications, Cloud Platform System, Helios, System Availability, Grafana, Backend, Servicebus, AI Platforms, Information Technology, Real Time Data, Data Management, Cloudwatch, Stream Processing, Data Pipelines, Jenkins, Microservices - **Published:** September 19, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=c32e8dd42a863c9f ## About the Role * Decision Making and Critical Thinking: Ability to lead the analysis and resolution of complex issues within distributed data platforms, designing scalable, and resilient solutions * Effective Communications: Ability to communicate across teams by sharing feedback constructively, listening to others, and creating documentation that makes data systems and processes easy to understand and support * Software Development: Experience in leading the design and development of backend systems and data pipelines using Python, Java, and modern frameworks, providing technical directions and ensuring the delivery of reliable, scalable solutions * Software Development Life Cycle: Experience leading the delivery of data engineering solutions in an Agile environment by guiding work through the full development lifecycle, translating requirements into technical solutions, and ensuring projects are delivered with quality, reliability, and business value * Software Integration Engineering: Capability to lead the design and integration of APIs, data pipelines, streaming platforms, and databases to enable reliable data exchange across enterprise systems and partner platforms * Software Product Design/Architecture: Expertise leading the design of scalable, event-driven data systems and architectures, guiding technical decisions and ensuring solutions are reliable, maintainable, and aligned with business needs. * Software Product Technical Knowledge: Ability to apply strong knowledge of AWS services and data engineering tools to define requirements, support testing and deployment activities, troubleshoot issues, and ensure data solutions are configured, implemented, and operated effectively across environments * Software Product Testing: Ability to define and implement testing strategies, including functional, performance, and data quality testing, to ensure reliable, scalable, and high-performing data solutions across the development lifecycle. Top Candidates Will Have: * Bachelor's degree in Computer Science, Computer Engineering, or related field * 8+ years of experience in data engineering or related disciplines with increasing responsibility * Extensive experience on modern, large scale, complex Caterpillar data platforms such as Helios Data Platform * Strong foundation developing and deploying Python solutions to a production environment * Experience leading teams to build high-throughput, scalable data pipelines * Strong hands-on experience with AWS data services (Kinesis, S3, DynamoDB, EventBridge, etc.) at scale * Strong in SQL, including data quality and validation practices * Experience in deploying software using CI/CD tools such as Azure DevOps, Jira, Jenkins, etc. * Experience developing microservices that support real-time data ingestion * Experience developing software applications using relational and noSQL databases * Ability to ensure data integrity across distributed and streaming systems * Experience with monitoring, testing, and automation in large-scale data environments ## Description As a Lead Data Engineer, you will design, build, and maintain scalable data pipelines, microservices, and cloud-based data platforms that deliver reliable, high-quality data for business and engineering teams. Working in an agile environment, you will help drive data architecture, performance, reliability, and continuous improvement across modern data solutions. What You Will Do: * Actively collaborate with Principal Software Engineers and Data Architects to define solution architecture * Lead the solution design and optimization of scalable data pipelines and microservices in Python, enabling both real-time and batch data processing across enterprise platforms * Drive the development of cloud-native data ingestion and streaming solutions leveraging AWS services including Kinesis, S3, DynamoDB, EventBridge, and related technologies * Own the design, implementation, and operational excellence of data integration frameworks and source data pipelines supporting CI Autonomy initiatives * Partner with business, product, and engineering stakeholders to translate complex requirements into scalable data architectures, workflows, mappings, and system designs * Establish and enforce automated testing, data quality controls, and validation frameworks to ensure integrity, reliability, and compliance across distributed data ecosystems * Lead operational monitoring, performance tuning, and root-cause analysis of production data platforms using observability tools such as CloudWatch to maintain high availability and service reliability ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Leveraging Real time data in FSIs](https://www.wearedevelopers.com/videos/806-leveraging-real-time-data-in-fsis) - [Improving quality with Agentic AI with Rovo Dev and Xray](https://www.wearedevelopers.com/videos/2005-improving-quality-with-agentic-ai-with-rovo-dev-and-xray) - [Modern Data Architectures need Software Engineering](https://www.wearedevelopers.com/videos/1030-modern-data-architectures-need-software-engineering) - [NoSQL Data Modeling for Front-end Developers](https://www.wearedevelopers.com/videos/297-nosql-data-modeling-for-front-end-developers) - [Collaboration Quantified: Lessons from Open Source Developer Networks](https://www.wearedevelopers.com/videos/1422-collaboration-quantified-lessons-from-open-source-developer-networks) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [The Most Popular IT Jobs on the Market](https://www.wearedevelopers.com/magazine/376-the-most-popular-it-jobs-on-the-market)