> Markdown version of [/jobs/ext/2620396-principal-data-engineer](https://www.wearedevelopers.com/jobs/ext/2620396-principal-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal Data Engineer - **Company:** CareerCircle - **Location:** San Diego, CA, United States (Remote available) - **Salary:** $124,800.0 - $208,000.0 - **Contract:** Temporary contract - **Skills:** Artificial Intelligence, Data Analysis, Cloud Computing, Cloud Database, Data Architecture, Information Engineering, Data Governance, Data Infrastructure, Data Security, Dataspaces, Data Systems, Digital Architecture, Distributed Computing Environment, Distributed Data Store, Distributed Systems, Python (Programming Language), Machine Learning, Meta-Data Management, DataOps, Enterprise Data Management, Data Processing, Cloud Platform System, High Performance Computing, Apache Spark, Data Strategy, Data Lakes, Information Technology, Data Management, Machine Learning Operations, Data Pipelines, Databricks - **Published:** August 1, 2026 - **Apply:** https://www.careercircle.com/jobs/all/all/usa/ca/san-diego/006ad480-c92a-4ae4-8bf9-f93ae66ac9dd ## About the Role Automation Mentorship Innovation Data Lakes Mathematics Scalability Reliability Data Quality Data Science Apache Spark Communication Observability Data Security Data Modeling Data Pipelines Access Controls Data Governance Data Processing Computer Science Azure Databricks Machine Learning Data Engineering Data Architecture Distributed Cloud Influencing Skills Technical Strategy Workflow Management Data Infrastructure Metadata Management Distributed Data Store Artificial Intelligence Engineering Design Process Python (Programming Language) Influencing Without Authority Continuous Improvement Process MLOps (Machine Learning Operations), * Strong hands-on development experience using Python to build, maintain, and optimize enterprise-scale data solutions. * Proven experience designing and implementing scalable data pipelines that integrate data from multiple scientific, engineering, and operational sources. * Expertise with Azure Databricks, Apache Spark, and distributed cloud-based data processing environments. * Deep understanding of data architecture, including data lakes, data modeling, metadata management, and platform scalability. * Experience establishing and managing data governance, including data quality, lineage, security, and access control standards within complex data environments. * Strong expertise in cloud-based data platforms and distributed systems, with practical experience in modern data technologies. * Experience designing and supporting large-scale, data-intensive engineering or scientific environments. * Experience working with research, simulation, instrumentation, sensor, or high-performance computing generated data. * Familiarity with AI and machine learning data pipelines, MLOps practices, or machine learning enablement. * Deep knowledge of data governance, security, metadata management, lineage, and data quality frameworks. * Ability to lead cross-functional initiatives and influence technical strategy across multiple stakeholder groups. * Strong communication skills with the ability to translate complex technical concepts into clear business value. * Bachelor's degree in Computer Science, Data Engineering, Data Science, Mathematics, or a related technical field. * 10+ years of experience in data engineering, data architecture, or enterprise data platform development., Research Visionary Automation Mentorship Innovation Data Lakes Mathematics Scalability Reliability Data Quality Data Science Apache Spark Communication Observability Data Security Data Modeling Data Pipelines Access Controls Data Governance Data Processing Computer Science Azure Databricks Machine Learning Data Engineering Data Architecture Distributed Cloud Influencing Skills Technical Strategy Workflow Management Data Infrastructure Metadata Management Distributed Data Store Artificial Intelligence Engineering Design Process Python (Programming Language) Influencing Without Authority Continuous Improvement Process MLOps (Machine Learning Operations) +0 ## Description The Principal Data Engineer leads the design, development, and evolution of an enterprise data architecture and data platform strategy that powers advanced research, simulation, testing, and engineering environments. This role builds and maintains scalable, cloud-based data pipelines and infrastructure that provide scientists, engineers, and machine learning teams with reliable, high-quality data for analytics, modeling, and technology development. The Principal Data Engineer establishes robust data governance standards, optimizes platform performance and cost, and serves as a hands-on technical leader who drives innovation, scalability, and continuous improvement across the data ecosystem while partnering closely with research, engineering, IT, and business stakeholders., * Lead the design, implementation, and roadmap of the enterprise data architecture and data platform strategy to support research, simulation, testing, and engineering environments. * Build, maintain, and optimize scalable data pipelines that integrate data from multiple scientific, engineering, and operational sources. * Develop and manage cloud-based data infrastructure that supports analytics, modeling, and technology development for scientists, engineers, and machine learning teams. * Establish and maintain data governance standards, including data security, lineage, metadata management, data quality, and lifecycle controls across complex data environments. * Optimize the performance, reliability, scalability, and cost efficiency of cloud-based data systems and distributed data processing platforms. * Partner with research, engineering, IT, and business stakeholders to align data architecture and platform capabilities with organizational objectives. * Support the integration of AI and machine learning workflows, high-performance computing environments, simulation platforms, and advanced research systems into the data platform. * Define and enforce architecture standards, best practices, and patterns for data engineering and data platform development. * Document architecture decisions, data models, and technical designs, and communicate complex concepts clearly to both technical and non-technical audiences. * Serve as a hands-on technical leader by contributing to code, reviewing implementations, and mentoring other engineers in data engineering best practices. * Drive continuous improvement initiatives across the data ecosystem, including automation, observability, and reliability enhancements. * Collaborate with cross-functional teams to design and support large-scale, data-intensive engineering and scientific environments. * Enable AI/ML data pipelines and MLOps practices that support machine learning experimentation, training, and deployment. * Ensure data platforms effectively support machine learning, simulation, modeling, experimental research, and high-performance computing workflows. * Influence technical strategy and lead cross-functional initiatives that shape the future of the research data ecosystem and data strategy. ## Related Videos - [Empowering Retail Through Applied Machine Learning](https://www.wearedevelopers.com/videos/976-empowering-retail-through-applied-machine-learning) - [Cutting LLM Costs Without Cutting Quality: How to Beat Proprietary LLMs with Fine-Tuned Open Source](https://www.wearedevelopers.com/videos/100151-cutting-llm-costs-without-cutting-quality-how-to-beat-proprietary-llms-with-fine-tuned-open-source) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Enabling intelligent logistics automation: home-grown Industrial IoT platform at Austrian Post](https://www.wearedevelopers.com/videos/2018-enabling-intelligent-logistics-automation-home-grown-industrial-iot-platform-at-austrian-post) - [OLTP in the Lakehouse: Redefining Data for AI Workloads](https://www.wearedevelopers.com/videos/2038-oltp-in-the-lakehouse-redefining-data-for-ai-workloads) - [The Data Mesh as the end of the Datalake as we know it](https://www.wearedevelopers.com/videos/156-the-data-mesh-as-the-end-of-the-datalake-as-we-know-it) ## Related Articles - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story)