> Markdown version of [/jobs/ext/3099493-data-scientist](https://www.wearedevelopers.com/jobs/ext/3099493-data-scientist). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Scientist - **Company:** Epitec, Inc. - **Location:** Austin, TX, United States - **Experience:** Expert - **Salary:** $201,760.0 - $223,392.0 - **Contract:** Temporary contract - **Skills:** Artificial Intelligence, Airflow, Cluster Analysis, Computer Engineering, Continuous Integration, Information Engineering, Data Infrastructure, Supervisory Control and Data Acquisition (SCADA), Python (Programming Language), Machine Learning, Operational Data Store, Power BI, Standard Sql, Azure Machine Learning, Statistical Process Control (SPC), Tableau (Software), Chatbots, Apache Spark, Fastapi, Pyspark, Information Technology, Data Analytics, Star Schema, Machine Learning Operations, Data Pipelines, Databricks - **Published:** September 26, 2026 - **Apply:** https://www.dice.com/job-detail/ce221852-c986-41bd-9b10-c3a33104d26b ## About the Role * Bachelor's or master's degree in a quantitative, technical, or engineering field such as Data Science, Computer Science, Computer Engineering, Statistics, Industrial/Mechanical/Manufacturing Engineering, Operations Research, Applied Mathematics, or a related field. * 8-10+ years of applied data science experience, including analysis, feature work, statistical or ML modeling on real operational or business datasets. * Strong experience with data modeling, data pipelines, and data analytics, particularly building something useful out of messy data. * Strong Python and SQL skills with evidence of working with large, messy tables, not only notebooks on clean extracts. * Experience producing models you can defend, including classification, regression, clustering, anomaly detection, or time series, with a clear target and validation approach. * Experience creating features from machine, sensor, process, quality, maintenance, or other operational data. * Ability to explore, clean, analyze, and derive meaningful findings from incomplete or fragmented datasets. * Ability to explain how a model was validated and determine when it is wrong or right. * Comfort telling stakeholders when a model should not ship. * Ability to learn an unfamiliar plant process quickly., * Experience in manufacturing, industrial IoT, semiconductor, automotive, aerospace, energy, telecom/operations, or equipment-heavy environments. * Manufacturing experience working with machine, sensor, process, quality, maintenance, or MES/historian data. * Experience with Databricks, Spark/PySpark, or similar cloud analytics. * Familiarity with MLOps concepts such as model tracking, monitoring, and drift while partnering with platform teams. * Experience with SPC, model explainability, historians, or MES data. * AI application/platform experience is a plus but not required. ## Description The team is building a brand-new AI application/system from the ground up and needs an experienced Data Scientist who can dig into manufacturing, quality, and operational data, build data models, and help identify new ways AI can support the business. Most of the job is in the data: understanding how a process behaves, cleaning noisy and incomplete signals, defining "normal" vs. "abnormal" with people who run the line, creating features, validating against real outcomes, and explaining limits when the data cannot support a model. You will not own the data platform. This is not a research lab role and not a platform-engineering role. This is an opportunity to help build an AI application from the ground up, work with extensive manufacturing and quality data, and contribute to solutions that can have a significant impact across manufacturing operations., * Frame manufacturing problems with plant and engineering partners; push back when labels, ground truth, or "accuracy" expectations are not real. * Explore, clean, and join fragmented operational data, including machines, sensors, quality, maintenance, production, and MES/historian extracts. * Dig into large amounts of manufacturing and quality data and help build the system from scratch. * Build and validate statistical and machine-learning models for anomaly, quality, health, and process monitoring. * Report false positives/negatives and business cost, not only a leaderboard metric. * Create features and define "normal" vs. "abnormal" behavior based on real operational data. * Hand usable outputs to engineers and operators, including thresholds, explanations, and "what to do when this fires." * Support models after they are in use. * Work with data engineering and software on pipelines, Databricks, and production. You are the customer of the platform, not the person hired to build it. * Review the available data, bring creative ideas to the team, and identify additional opportunities to help the business with AI. * Work on problems such as process drift, abnormal machine behavior, quality prediction, equipment health, bottlenecks, downtime, scrap/rework, and root-cause support., * Data engineering/data modeling centered on lakehouse, medallion, star schema, Unity Catalog, ADF, or building "pipelines for the DS team" * MLOps/ML platform work such as Airflow, SageMaker plumbing, FastAPI services, or CI/CD with little hands-on dataset analysis * GenAI product work such as RAG chatbots, LangChain agents, or Copilot applications as the primary experience * BI/reporting work centered on Power BI/Tableau KPI applications without model development * Simply listing MES, SCADA, PLC, JPH, Databricks, or anomaly detection without being able to explain the dataset, finding, model, and validation result ## Related Videos - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [Fully Orchestrating Databricks from Airflow](https://www.wearedevelopers.com/videos/336-fully-orchestrating-databricks-from-airflow) - [Build your backend using FastAPI](https://www.wearedevelopers.com/videos/506-build-your-backend-using-fastapi) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Empowering Retail Through Applied Machine Learning](https://www.wearedevelopers.com/videos/976-empowering-retail-through-applied-machine-learning) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [The Most Popular IT Jobs on the Market](https://www.wearedevelopers.com/magazine/376-the-most-popular-it-jobs-on-the-market) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production)