> Markdown version of [/jobs/ext/3607552-lead-data-engineer](https://www.wearedevelopers.com/jobs/ext/3607552-lead-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Lead Data Engineer - **Company:** Techridge, Inc. - **Location:** Cary, NC, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** LangGraph Framework, AI Evaluation, Microsoft Azure, Continuous Integration, Information Engineering, Data Infrastructure, Data Profiling, Graph Database, Information Retrieval, Python (Programming Language), Metadata Repositories, Neo4j, Performance Tuning, Posix, Prism (Software), Role-Based Access Control, Regression Testing, OpenAI, Standard Sql, Azure Data Lake, Scala (Programming Language), SPARQL, SQL Databases, Data Streaming, Management of Software Versions, Azure Data Factory, Great Expectations (Foster Youth College-readiness and Support Program), Retrieval-Augmented Generation, Large Language Models, Prompt Engineering, Apache Spark, Llamaindex, Data Lakes, Pyspark, Integration Frameworks, Machine Learning Operations, Terraform, Data Pipelines, Human in the Loop, Workday, Key Vault, Databricks - **Published:** October 7, 2026 - **Apply:** https://www.dice.com/job-detail/990db7c5-599f-4302-bfa6-2e8b219c2468 ## About the Role * 12-18 years of overall experience in Data Engineering / Data Platform delivery. * Expert-level Python, Scala, and PySpark with production-ready, modular, and well-tested solutions. * Strong experience troubleshooting Spark workloads and optimizing large-scale batch and streaming pipelines using Delta Lake. * Strong SQL and Data Modeling skills, including dimensional/normalized modeling, schema design, and data contracts. * Deep hands-on expertise with Databricks, including Delta Lake, Unity Catalog, Jobs & Workflows, cluster/pool management, performance tuning, and Model Serving. * Strong Azure Data experience with ADLS Gen2, Azure Data Factory, and Azure Event Hubs. * 3+ years of production experience designing and delivering LLM-based solutions, including RAG, agentic/tool-calling workflows, chunking, embeddings, vector/hybrid retrieval, and prompt engineering. * Hands-on experience with LangChain, LlamaIndex, or LangGraph and at least one provider stack such as Azure OpenAI, OpenAI, or Databricks Model Serving. * Strong AI evaluation practices including golden datasets, regression testing, accuracy/hallucination tracking, and human-in-the-loop feedback. * Experience building metadata-driven data frameworks, including schema inference, data profiling, lineage, and data catalogs. * Proven experience delivering enterprise-scale Medallion / Lakehouse architectures. * Strong Azure security and governance knowledge: Entra ID, Managed Identities, RBAC, POSIX ACLs, Key Vault, Private Endpoints, and PII handling. * Experience with CI/CD and Infrastructure as Code, including Azure DevOps, Terraform, Databricks Asset Bundles, and automated data pipeline testing. * Excellent technical communication skills with the ability to document and present architecture decisions to both technical and non-technical stakeholders. Strongly Preferred * Knowledge graphs and ontologies: RDF/SPARQL, Neo4j, and graph modeling. * Enterprise-scale Text-to-SQL or semantic-layer-based natural language query systems. * ML-based anomaly detection for time-series or transactional financial data. * Experience in Financial Services or Insurance, including financial close, GL, subledger, reconciliation, or actuarial data. * LLMOps / MLOps experience including model/prompt versioning, cost governance, and observability. * Certifications such as Databricks Data Engineer Professional, Azure DP-203/DP-700, or AZ-305. * Experience with dbt, Great Expectations, Workday, Prism, or Accounting Center.