> Markdown version of [/jobs/ext/1425674-principal-data-engineer](https://www.wearedevelopers.com/jobs/ext/1425674-principal-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal Data Engineer - **Company:** Realtime Software Solutions, LLC - **Location:** United States (Remote available) - **Experience:** Expert - **Salary:** $155,000.0 - $195,000.0 - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Microsoft Azure, Clinical Data Repository, Cloud Computing, Data Architecture, Data Discovery, Data Governance, Data Visualization, Relational Databases, Database Queries, Interoperability, Python (Programming Language), Machine Learning, Natural Language Processing, Named Entity Recognition, Power BI, Azure Machine Learning, SQL Databases, Management of Software Versions, Large Language Models, Prompt Engineering, Technical Debt, Generative AI, Pandas, Scikit Learn, Uipath, Information Technology, Deployment Automation, HuggingFace, Data Analytics, Performance Monitor, Machine Learning Operations, Tools for Reporting, Spacy, Crosswalk, Data Pipelines - **Published:** July 24, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=59d7d8c485e8269d ## About the Role * Bachelor's degree in data science, Computer Science, Mathematics, Statistics, Economics, or a related quantitative field, or equivalent professional experience. * 5+ years of experience in data science, data analytics, or a related discipline, including production ML/AI deployments. * Strong proficiency in Python for data science workflows, including pandas, scikit-learn, and NLP libraries (e.g., spaCy, Hugging Face Transformers). * Proven experience designing and delivering NLP pipelines and/or forecasting models in a business context. * Solid command of SQL for data querying, transformation, and analysis across relational databases. * Experience with BI and reporting tools, particularly Power BI, including data modeling and DAX. * Demonstrated ability to communicate analytical findings clearly to non-technical stakeholders and drive decision-making. * Experience working in regulated industries (healthcare, finance, or similar) with an understanding of compliance and data governance requirements. WHAT SETS YOU APART? * 7+ years of data science or analytics experience, including a team lead or principal contributor role. * Experience with Large Language Models (LLMs), Retrieval-Augmented Generation (RAG), and prompt engineering for enterprise use cases. * Familiarity with cloud-based ML platforms (GCP, AWS, or Azure) and MLOps practices. * Experience with process automation tools (e.g., UiPath or similar RPA platforms). * Working knowledge of process optimization frameworks such as Lean Six Sigma (Green Belt or higher). * Exposure to clinical data standards, health data interoperability, or cross-client data standardization projects. * Proficiency in data visualization and dashboard design; PL-300 Power BI Data Analyst certification is a plus. ## Description The Principal Data Engineer serves as the technical and analytical authority for the organization's data science practice. This is a senior individual contributor and team lead role responsible for driving AI/ML strategy, delivering data-driven solutions, and elevating the analytical capability of the wider team. The role combines deep technical expertise in machine learning, NLP, and forecasting with strong product ownership and stakeholder communication skills, ensuring that data science investments translate into measurable business outcomes., * Define and drive the data science and AI roadmap, aligning model development priorities with business objectives and product strategy. * Lead end-to-end delivery of ML and AI solutions - from problem framing, data discovery, and model design through validation, deployment, and performance monitoring. * Translate ambiguous business problems into well-scoped data science workstreams, identifying quick wins alongside longer-term strategic initiatives. * Champion best practices in model development, including versioning, documentation, validation, and observability. 2. NLP, Forecasting & Advanced Analytics * Design and implement NLP pipelines for use cases such as entity extraction, semantic mapping, classification, and retrieval-augmented generation (RAG). * Build and maintain forecasting and predictive models to support operational and strategic decision-making. * Apply statistical and machine learning methods to identify root causes of process inefficiencies and data quality issues. * Develop reusable data pipelines, crosswalk tables, and transformation workflows that support scalable, cross-functional data products. 3. Data & Process Analysis * Conduct current-state assessments of data architecture, sources, and quality; define future-state data models and governance standards. * Develop and maintain KPI reporting frameworks and dashboards that enable performance monitoring and data-driven decision-making. * Apply process optimization methodologies (e.g., Lean Six Sigma) to identify bottlenecks, reduce cycle time, and improve data accuracy. * Ensure analytical outputs are accurate, auditable, and aligned with regulatory and compliance requirements (e.g., HIPAA, GDPR). 4. Stakeholder Collaboration & Communication * Partner closely with Product, Engineering, and business stakeholders to clarify requirements, validate feasibility, and define measurable success criteria. * Communicate complex analytical findings and model outputs clearly to both technical and non-technical audiences, including executive stakeholders. * Define value-realization strategies for data and AI investments, ensuring ROI is tracked through improved search, reporting, and operational insight. 5. Team Development & Knowledge Transfer * Mentor data analysts and junior data scientists through pairing, design reviews, and structured technical guidance. * Lead knowledge transfer of owned models, pipelines, and analytical frameworks to ensure team resilience and continuity. * Drive a culture of continuous learning, analytical rigor, and responsible AI within the data science function. Performance at this level is evaluated across four dimensions: Technical Quality * Models, pipelines, and analytical frameworks consistently meet peer review standards and produce reliable, reproducible results. * Data quality, model drift, and technical debt in areas of ownership trend downward over time. Delivery & Impact * Data science deliverables are completed on schedule with accurate effort estimation; risks and blockers are surfaced early. * Analytical insights are actionable, with measurable improvements in KPIs such as process efficiency, data accuracy, or cost reduction. Organizational Impact * Data analysts and engineers who regularly collaborate with this role demonstrably improve analytical judgment and technical practice. * Data science recommendations are trusted by peers, product leadership, and executive stakeholders without requiring repeated validation. Communication & Leadership * Ambiguous analytical problems are framed and decomposed independently, without waiting for direction. * Findings and model outputs are presented in a way that drives clear decisions by non-technical stakeholders. ## Related Videos - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [How to achieve web automation with UiPath](https://www.wearedevelopers.com/videos/310-how-to-achieve-web-automation-with-uipath) - [Advanced Typing in TypeScript](https://www.wearedevelopers.com/videos/496-advanced-typing-in-typescript) - [Data Science on Software Data](https://www.wearedevelopers.com/videos/162-data-science-on-software-data) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [RPA crash course for .Net developers – intro into the world of RPA from the perspective of a .Net developer](https://www.wearedevelopers.com/videos/271-rpa-crash-course-for-net-developers-intro-into-the-world-of-rpa-from-the-perspective-of-a-net-developer) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Data Analyst Salary in the UK](https://www.wearedevelopers.com/magazine/278-data-analyst-salary-in-the-uk)