> Markdown version of [/jobs/ext/2548322-lead-data-engineer](https://www.wearedevelopers.com/jobs/ext/2548322-lead-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Lead Data Engineer - **Company:** ACCLINATE LLC - **Location:** United States (Remote available) - **Experience:** Expert - **Salary:** $79,000.0 - $116,000.0 - **Contract:** Permanent contract - **Skills:** Query Performance, Sql Data Warehouse, Data Analysis, Automation of Tests, BigQuery, Code Review, Customer Data Management, Data Validation, Data Cleansing, Data Governance, Data Infrastructure, Data Integration, Extract Transform Load (ETL), Data Mart, Data Retrieval, Data Security, Data Sharing, Dataspaces, Data Systems, Data Warehousing, Google Analytics, Identity and Access Management, Machine Learning, Meta-Data Management, Operational Databases, DataOps, SQL Databases, Data Streaming, Tokenization, Business Intelligence Development Studio, Google Cloud, Data Storage Technologies, Feature Engineering, Data Ingestion, Large Language Models, Snowflake, Database Optimization, Data Strategy, Git, Event Driven Architecture, Data Lakes, Data Lineage, Machine Learning Operations, Hubspot, Software Version Control, Data Pipelines, Databricks - **Published:** August 9, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=7cb5b9090b48377f ## About the Role * 6+ years building and operating production data pipelines, including at least 2 years owning a data platform end to end. * Demonstrated ownership of a cloud data warehouse in production. BigQuery preferred; Snowflake, Redshift, or Databricks considered. * Experience handling regulated health data (HIPAA/PHI) or comparable regulated data. * Bachelor's degree in a technical field, or equivalent practical experience. Required Tools and Technical Skills * SQL and BigQuery (primary data platform); Google Cloud Platform (IAM, storage, jobs) * Python for data preparation, imputation, and analysis notebooks * Pipeline and automation tooling (e.g., n8n or equivalent) and scheduled job management * HubSpot CRM data model and its BigQuery integration (Operations Hub Enterprise) * BI tooling, particularly Metabase * ML/MLOps exposure: VertexAI and feature pipelines (supporting production models, not research-grade model development) * Health-data privacy and compliance awareness (HIPAA/PHI) * Practical AI tooling experience, including building and validating LLM-based "skills"/agents Nice to Haves * HubSpot Ops Hub data model and its BigQuery sync; VertexAI and production ML feature pipelines; building and validating LLM-based agents or "skills". Strong candidates on the required list will be considered without these. ## Description As Acclinate's Lead Data Engineer, you will design, build, and operate the data infrastructure that powers the company's analytics, CRM, marketing, and machine learning initiatives. This is a hands-on role: you will be equally responsible for long-term architectural strategy and for the day-to-day engineering, pipeline maintenance, and reporting work that keeps Acclinate's data ecosystem running. You'll work across our event-driven architecture, CRM data platform, BI tooling, and production ML systems, partnering closely with the data lead, engineering, marketing, community, and customer success teams., * Design Event-Driven Data Models and Tables: Architect logical and physical data models, creating necessary data tables and schemas to effectively capture, store, and make accessible vast amounts of data generated by event streams. This includes defining how events are transformed into structured data suitable for analysis. * Orchestrate Data Integration for Advanced Analytics: Design integration strategies to combine data from event-driven systems with existing data sources, ensuring a unified and consistent view for sophisticated analyses, including data warehousing and data mart design. * Co-develop Acclinate's Data Strategy: Collaborate with appropriate unit leads and team members to define Acclinate's long-term vision for data, aligning it with business objectives, and creating a roadmap for implementation across both traditional and event-driven data. * Contribute to Data Governance Frameworks: Collaborate with appropriate unit leads and team members to define data ownership, stewardship, quality standards, metadata management, and data lineage tracking for all data, including event data. * Ensure Data Security and Compliance: Partner with engineering and legal teams to implement data security controls, access controls, and ensure adherence to relevant data privacy regulations (including PII and sensitive health information) for all data types. * Optimize Data Storage and Performance: Recommend and implement appropriate data storage technologies and indexing strategies to ensure efficient data retrieval and analytical query performance, especially for high-volume event data. * Evaluate and Recommend New Technologies: Stay abreast of industry trends and propose new data technologies and tools that can benefit Acclinate's data platform. Hands-On Data Operations & Execution: * Oversee Data Integration and ETL/ELT Processes: Build and maintain efficient, reliable data pipelines (e.g., in n8n or equivalent tooling) that transform event streams and other source data into a unified data environment such as a data warehouse or data lake, primarily on GCP/BigQuery. * Own CRM Data (HubSpot): Manage HubSpot property and schema design, maintain the HubSpot-to-BigQuery integration (Operations Hub Enterprise), and perform ongoing data cleanup and reconciliation. * Build and Maintain BI/Analytics Enablement: Build and maintain Metabase dashboards and a shared metrics library; produce recurring monthly and quarterly reports for Customer Success, Community, Marketing, and project reporting needs. * Support Production Machine Learning (PPI Model): Own the ongoing data needs of a live production ML model, including feature engineering, data preparation and imputation, and managing model data in VertexAI; coordinate deployment-related data inputs with the product and engineering teams. * Manage Marketing, Social, and Survey Data Pipelines: Build and maintain daily ingestion pipelines from Meta, Instagram, YouTube, Google Analytics, and Typeform into BigQuery. * Manage External Data Partnerships: Own tokenized- and identified-data workflows and vendor relationships to support secure, compliant data sharing. * Build AI Tooling: Develop and maintain Claude data skills for non-technical teams, including appropriate validation and guardrails. * Maintain Data Documentation and Handover Materials: Create and maintain data dictionaries, runbooks, and onboarding materials to support knowledge transfer and business continuity. * Apply Engineering Best Practices: Manage all pipeline and transformation code in version control (git), participate in code review, and maintain automated tests and data quality checks on business-critical pipelines. * Own Data Quality and Reliability: Define and monitor freshness, completeness, and accuracy checks on critical tables; triage pipeline failures and data incidents, and communicate impact and resolution to affected teams. * Manage Platform Cost: Monitor and optimize BigQuery and GCP spend, including query cost, storage tiering, and scheduled job efficiency. * Support Compliance Operations: Maintain least-privilege access controls and periodic access reviews on data assets, apply Acclinate's de-identification standard to externally shared datasets, and produce evidence artifacts (access logs, data lineage, data flow documentation) for audits and client security reviews. ## Related Videos - [Integrate your Cognitive Assistant with 3rd-party DBs and software](https://www.wearedevelopers.com/videos/249-integrate-your-cognitive-assistant-with-3rd-party-dbs-and-software) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Git for Code Reviews](https://www.wearedevelopers.com/videos/429-git-for-code-reviews) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [Stop Committing Your Secrets - GIt Hooks To The Rescue!](https://www.wearedevelopers.com/videos/573-stop-committing-your-secrets-git-hooks-to-the-rescue) ## Related Articles - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again)