> Markdown version of [/jobs/ext/1212997-lead-data-engineer](https://www.wearedevelopers.com/jobs/ext/1212997-lead-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Lead Data Engineer - **Company:** Datasage Technologies - **Location:** Mountain View, CA, United States - **Experience:** Expert - **Salary:** $93,600.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Airflow, Amazon S3, JIRA, Information Engineering, Data Governance, Extract Transform Load (ETL), Data Mart, Data Warehousing, Cursor (Graphical User Interface Elements), Python (Programming Language), Productivity Software, SQL Databases, Large Language Models, Apache Spark, Generative AI, Data Layers, Pyspark, Graphql, Splunk, Data Pipelines, Amazon Elastic Mapreduce (EMR), Pagerduty, Databricks - **Published:** July 9, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=49eb98bc6ec961a8 ## About the Role 8+ years in Data Engineering, with strong hands-on Spark (PySpark/Scala), SQL, and Python. * Experience building and operating ETL/ELT pipelines on cloud platforms (AWS EMR/S3, Databricks, or equivalent); workflow orchestration (Airflow or similar). * Comfortable owning pipeline operations end-to-end: reading dependency graphs, diagnosing failures, working with on-call /PagerDuty/Splunk, and driving fixes with multiple teams. * Working knowledge of data warehousing/data mart design, data governance, and PII handling practices. * AI-native mindset: familiarity with LLM capabilities, evaluation frameworks, and the creative application of AI principles to engineering challenges. Proficiency with GenAI productivity tools (e.g., Claude, Cursor, Codex) to enhance engineering workflows. * Strong communicator; comfortable directing an offshore IDC team with minimal handholding. ## Description We're looking for a hands-on Data Engineer who can operate independently. High-visibility, high-ownership role at the center of T4I's HR data platform direct exposure to AI initiatives, cross-functional stakeholders, and the chance to meaningfully reduce day-to-day operational load for the team. What You'll Do * Build, maintain, and troubleshoot Spark/EMR ETL pipelines feeding HR and Workforce data marts. * Monitor and remediate Data Asset Score issues (data quality rules, governance/PII-minimization actions) to keep HR data assets compliant ahead of deadlines. * Be the daily coordination point between US HR business stakeholders, the IDC (India) engineering team, and platform/infra teams - translating requirements, unblocking IDC, and reporting status. * Triage and resolve Jira tickets/bugs raised against HR data mart pipelines; write clear runbooks and pipeline documentation. * Build semantic layers and Retrieval-Augmented Generation (RAG) pipelines. * Integrate REST and/or GraphQL APIs into data workflows. * Operate with minimal oversight: escalating only true blockers and proactively flag risks before they become incidents. What You Bring ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [Collaboration Quantified: Lessons from Open Source Developer Networks](https://www.wearedevelopers.com/videos/1422-collaboration-quantified-lessons-from-open-source-developer-networks) - [Enjoying SQL data pipelines with dbt](https://www.wearedevelopers.com/videos/823-enjoying-sql-data-pipelines-with-dbt) - [AI Model Management Life Circles: ML Ops For Generative AI Models From Research to Deployment](https://www.wearedevelopers.com/videos/1152-ai-model-management-life-circles-ml-ops-for-generative-ai-models-from-research-to-deployment) - [Integrate your Cognitive Assistant with 3rd-party DBs and software](https://www.wearedevelopers.com/videos/249-integrate-your-cognitive-assistant-with-3rd-party-dbs-and-software) ## Related Articles - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production)