Data Architect 5- Data Scientist

eShocan LLC
Austin, TX, United States
6 days ago
Apply on www.careerjet.com
Prepare application

Role details

Contract type
Temporary to permanent
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
6 years minimum
Compensation
$160,000.0 - $180,000.0
Working hours
Regular working hours

Tech stack

Airflow Continuous Integration Data Architecture Information Engineering Data Infrastructure Supervisory Control and Data Acquisition (SCADA) Python (Programming Language) Machine Learning Operational Data Store Power BI Azure Machine Learning Software Engineering
+11 more
Statistical Process Control (SPC) SQL Databases Tableau (Software) Data Processing Chatbots Apache Spark Fastapi Pyspark Star Schema Machine Learning Operations Databricks

Job description

We are hiring a Senior Data Scientist to work with factory and industrial data: find what is going wrong (or about to), build models that hold up against plant reality, and help the team act on them. Most of the job is in the data - understanding how a process behaves, cleaning noisy and incomplete signals, defining “normal” vs “abnormal” with people who run the line, creating features, validating against real outcomes, and explaining limits when the data cannot support a model. You will partner with manufacturing, quality, maintenance, data engineering, and software. You will not own the data platform. This is not a research lab role and not a platform-engineering role. What we are hiring for Someone who can walk a real example: this was the grain of the data, this is what I found, this is the model, this is how I knew it was wrong or right, this is what operations did with it. Typical problems: process drift, abnormal machine behavior, quality prediction, equipment health, bottlenecks, downtime, scrap/rework, root-cause support. Methods follow the problem (statistical limits, clustering, isolation forest, time series, autoencoders, supervised models when labels exist) - we do not hire to a method list. Manufacturing experience is a plus. We will also consider people from industrial IoT, equipment, quality, automotive, semiconductor, energy, telecom/ops, or similar operational environments who have done this loop on messy sensor or process data. Responsibilities Frame manufacturing problems with plant and engineering partners; push back when labels, ground truth, or “accuracy” expectations are not real. Explore, clean, and join fragmented operational data (machines, sensors, quality, maintenance, production, MES/historian extracts - you do not need to have used every acronym). Build and validate statistical and machine-learning models for anomaly, quality, health, and process monitoring; report false positives/negatives and business cost, not only a leaderboard metric. Hand usable outputs to engineers and operators (thresholds, explanations, “what to do when this fires”), and support models after they are in use. Work with data engineering and software on pipelines, Databricks, and production - you are the customer of the platform, not the person hired to build it., Data engineering / data modeling (lakehouse, medallion, star schema, Unity Catalog, ADF, “pipelines for the DS team”) MLOps / ML platform (Airflow, SageMaker plumbing, FastAPI services, CI/CD) with little analysis of a dataset GenAI product work (RAG chatbots, LangChain agents, Copilot apps) as the primary story BI/reporting (Power BI/Tableau KPI apps) without model work Resumes that list MES, SCADA, PLC, JPH, Databricks, and anomaly detection but never name a dataset, a finding, or a validation result Those skills exist on the team or in partner teams. We need the person who works the data. No immigration related sponsorship will be provided for this role. Please do not apply for this role if you require employer sponsorship. This includes direct company sponsorship or entry of an employer as the immigration employer of record or any work authorization requiring written immigration support from an employer. All candidates need to be legally eligible to work in the United States.

Requirements

Bachelor’s or master’s in a quantitative or engineering field (data science, CS, statistics, industrial/mechanical/manufacturing engineering, OR, applied math, or related). 6+ years of applied data science (analysis, feature work, statistical or ML modeling on real operational or business datasets). Count data-science years, not total years in IT, DBA, or software engineering. Strong Python and SQL; evidence of working large, messy tables - not only notebooks on clean extracts. Production of models you can defend: classification, regression, clustering, anomaly detection, or time series, with a clear target and validation approach. Experience creating features from machine, sensor, process, quality, maintenance, or other operational data (industrial preferred; high-volume ops data from adjacent domains is acceptable). Comfort telling stakeholders when a model should not ship. Ability to learn an unfamiliar plant process quickly. Preferred qualifications Time in manufacturing, industrial IoT, semiconductor, automotive, aerospace, energy, or equipment-heavy operations. Databricks, Spark/PySpark, or similar cloud analytics (we use Databricks; we do not require you to have been the lakehouse owner). Familiarity with MLOps (tracking, monitoring, drift) as a partner to platform teams. SPC, explainability, or prior work with historians/MES data.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerjet.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

2:19 min

Introduction to Apache Airflow for advanced orchestration

Alan Mazankiewicz · LIVE

1:22 min

Key features and advantages of using FastAPI

Ashmi Banerjee · World Congress 2022

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

3:37 min

Scaling machine learning pipelines from prototypes to petabytes

Julian Joseph · LIVE

1:34 min

Bringing diverse skills to industrial data science roles

Katja Träumner

Videos

See all

Related articles

See all