> Markdown version of [/jobs/ext/2435922-founding-data-engineer](https://www.wearedevelopers.com/jobs/ext/2435922-founding-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Founding Data Engineer - **Company:** Percepta LLC - **Location:** New York, United States - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Information Engineering, Machine Learning, Operational Databases, Large Language Models, Data Management, Databricks - **Published:** August 7, 2026 - **Apply:** https://www.dice.com/job-detail/37de9f2c-2eae-4fc5-9eec-62c6ee0b8503 ## About the Role You might come from any point on the spectrum - a strong data engineer; a software engineer who's done real data work; someone who's done data science and software; or an ML engineer who now wants to build more. What's common: you can build in ambiguity, you form opinions and ship, and you care about building leverage, not just outputs. * Strong experience around some combination of Data Science, Data Engineering, Machine Learning. * A product instinct for the second half of the job - you want to build the thing that makes the work easier, not just do the work * Intuition for what modern AI/ML and LLM systems actually need from data (features, retrieval, context, embeddings) * High ownership and strong communication - you're comfortable embedded directly with customer teams Nice To Have * Experience building agentic or automated data-engineering tooling * Hands-on experience with modern cloud data platforms (e.g., Databricks) * Experience with health-system data (EHR, claims, and other operational healthcare datasets) or other complex, regulated enterprise data * Prior startup, founding, or forward-deployed experience ## Description We're hiring one of the founding members of Percepta's data team - a role that lives across the full spectrum from data engineering to data science to ML engineering. You won't be boxed into one of those; the best person here has a center of gravity in one and real range across the others. The job has two halves, and you'll do both: 1. Be the data person. Build the pipelines, models, analysis, "data packs," and ontology that turn messy enterprise data into something AI can actually use - and do it fast, inside real customer environments. 2. Build the product around that. Build the tooling, abstractions, and increasingly agentic/automated systems that make the first half faster and compounding across every customer we work with. This is where you set the taste and help form our strategy for how Percepta does data - not as a one-off, but as something that gets better every time we do it. As a founding hire, you're not inheriting a playbook - you're writing it. What You'll Do * Build end-to-end pipelines and models that turn fragmented, messy enterprise data into high-leverage, AI-ready assets * Structure and normalize noisy datasets - defining the data packs and ontology that our AI engineers build on top of * Build the internal product and tooling that makes data work faster and repeatable across customers, so each engagement compounds rather than starts from zero * Work directly with operators and product/AI engineers to turn high-value use cases into production data workflows * Form strong technical opinions on data models, storage, orchestration, and infra tradeoffs - and make the calls ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Introduction to Azure Machine Learning](https://www.wearedevelopers.com/videos/368-introduction-to-azure-machine-learning) - [It's all about the Data](https://www.wearedevelopers.com/videos/425-it-s-all-about-the-data) - [Cutting LLM Costs Without Cutting Quality: How to Beat Proprietary LLMs with Fine-Tuned Open Source](https://www.wearedevelopers.com/videos/100151-cutting-llm-costs-without-cutting-quality-how-to-beat-proprietary-llms-with-fine-tuned-open-source) - [OLTP in the Lakehouse: Redefining Data for AI Workloads](https://www.wearedevelopers.com/videos/2038-oltp-in-the-lakehouse-redefining-data-for-ai-workloads) - [Modern Data Architectures need Software Engineering](https://www.wearedevelopers.com/videos/1030-modern-data-architectures-need-software-engineering) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development)