Data Scientist

TMC
United States
6 days ago
Apply on arc.dev
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Languages
English, German
Job source

Tech stack

Agile Methodology Artificial Intelligence Data Analysis Big Data Databases Data Cleansing Data Structures Python (Programming Language) Machine Learning NumPy Scrum Methodology Azure Machine Learning
+15 more
Supervised Learning Data Processing Cloud Platform System Feature Engineering Model Validation Data Strategy Pandas Matplotlib Scikit Learn Information Technology Data Analytics Xgboost Data Management Data Pipelines Databricks

Job description

As part of the Computational & Data Science team, you’ll work at the crossroads of science, engineering, product development, and business, turning experimental, product, and process data into powerful analytical solutions that enable smarter decision-making.

Our client - A global leader with a strong commitment to power the vehicles of the future hy accelerating the transition to clean mobility & developing breakthrough materials technology.

You - A hands-on (> mid) Data Scientist passionate about AI + Machine Learning eager to leverage statistical modelling and advanced predictive analytics to support faster product development cycles, improved end-to-end model performance, and generate value across the R&D ecosystem. Experienced in iterative delivery environments - Agile, Scrum, Kanban.

Hanau Full-time Hybrid (3 days onsite/week) 12-month assignment Occasional travel

How you’ll make a Difference:

  • Partner with stakeholders to turn business needs and scientific questions into data-driven strategy

  • Design, develop, and enhance ML models for catalyst performance prediction

  • Apply advanced analytics techniques, exploratory data analysis, feature engineering, model selection, hyperparameter tuning, and model performance assessment

  • Document model assumptions, methods, validation results, and recommendations in a reproducible way

  • Explore relevant databases and available data sources to understand data structures, completeness, quality, and usability for modelling

  • Unlock insights from correlations, patterns, outliers, and potential performance drivers within catalyst, emissions, laboratory, product, and process data

  • Develop scalable workflows that support efficient re-use of cleaned datasets

  • Work with catalyst, process, product, and laboratory data to identify performance drivers and relevant modelling features

  • Help advance the transition from exploratory analysis and prototype models to robust, usable modelling assets for the Project Team

  • Communicate insights and recommendations in a clear, compelling way that inspires action across technical and non-technical stakeholders

Requirements

  • Master’s degree in Data Science/Statistics/Mathematics/Computer Science/Engineering/Physics, Chemistry/Materials Science or a related discipline

  • Around 5 years of expertise in data science, machine learning, statistical modelling, or predictive analytics

  • Solid Python programming skills and hands-on experience with pandas, NumPy, scikit-learn, matplotlib

  • Proven expertise with database exploration, data cleaning/preparation, exploratory data analysis, and correlation analysis

  • Experience in developing and improving predictive models using complex technical or scientific datasets, in working with structured data from industrial, engineering, manufacturing, laboratory, or R&D environments

  • Ability to translate analytical findings into clear visualizations, technical documentation, and model interpretation materials

What will set you apart

  • Exposure to chemical engineering, catalyst development, emissions systems, materials science, automotive R&D, or related industrial research environments

  • Experience with supervised learning, ensemble methods, gradient boosting, regression/classification models, Bayesian modelling, hybrid modelling, or physics-informed machine learning

  • Knowledge of SQL databases, data pipelines, large datasets, or industrial data platforms

  • Familiarity with Azure Machine Learning, Databricks, cloud-based data platforms, or comparable

  • Experience with experimental design, laboratory data, process data, or product development data

  • A strong foundation in machine learning together with the ability to collaborate closely with domain experts to connect scientific parameters and engineering constraints to ML features and model outputs

Non-Technical

  • Effective in cross-functional and international R&D environments with the ability to take ownership of assigned analysis and modelling workstreams

  • Confident in explaining findings, correlations, model outcomes, and limitations to technical and non-technical audiences

  • Strong analytical mindset with the ability to connect data, domain knowledge, and business objectives

  • English fluent spoken and written

  • German nice to have.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on arc.dev
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:14 min

Structuring career paths and localized data architectures

Ulrich Wurstbauer +1 · LIVE

1:09 min

Configuring synthetic data for safe interactive programming

Mingshen Sun Mingshen Sun · World Congress 2024

2:03 min

Accelerating pandas dataframes using cudf module plugins

Ankit Patel Ankit Patel · World Congress 2024

2:34 min

Maximizing execution memory effectively via python numpy broadcasting

Jodie Burchell · LIVE

9:47 min

Transforming tabular metrics into meaningful business value dashboards

Boris Krumrey +2 · LIVE

6:58 min

Analyzing production code coverage data using pandas

Markus Harrer Markus Harrer · World Congress 2021

Videos

See all

Related articles

See all