Data Research Engineer

Fundamental
Laza, Spain
2 days ago
Apply on www.buscojobs.com.es
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours

Tech stack

Artificial Intelligence Extract Transform Load (ETL) Distributed Computing Environment Python (Programming Language) Machine Learning Cisco Nexus Switches NumPy Open Source Technology Software Engineering Data Processing Data Storage Technologies Delivery Pipeline
+7 more
Large Language Models Apache Spark Pandas Information Technology Dask Machine Learning Operations Data Generation

Job description

About FundamentalFundamental is an AI company pioneering the future of enterprise decision-making. Founded by DeepMind alumni, Fundamental has developed NEXUS - the world’s most powerful Large Tabular Model (LTM) - purpose-built for the structured records that actually drive enterprise decisions. Backed by world class investors and trusted by Fortune 100 companies, Fundamental unlocks trillions of dollars of value by giving businesses the Power to Predict.At Fundamental, you’ll work on unprecedented technical challenges in foundation model development and build technology that transforms how the world’s largest companies make decisions. This is your opportunity to be part of a category-defining company from the ground-up. Join the team defining the future of enterprise AI.Key responsibilitiesThe greatest research is done through solid engineering. As part of the research team, you will contribute to development of breakthrough machine learning models by working on one of the most crucial aspects of ML model training:data. The main responsibilities of this role are:Helping to identify, characterize and evaluate data sources, including realistic synthetic data generated from Structured Causal Models and physical / systems-based simulatorsBuilding and maintaining ETL pipelinesDesigning and implementing scalable, reliable data storage solutionsCollaborating with the rest of the research team to maintain a reliable, efficient training pipeline where data is a critical componentCollaborating with the wider engineering and infrastructure teamMust haveExperience with:Identifying good data sources to train and evaluate ML models, including real-world and realistic synthetic data sourcesBringing data from structured and unstructured sources, as well as simulators and causal models, into formats accessible by ML modelsStrong fundamentals of software engineeringStrong knowledge of:PythonPython data processing stack (numpy, pandas, …)Familiarity with:distributed processing (e.g. Ray, Dask Spark, Beam)data storage solutionsBasic ML knowledgeNice to haveContributions to open source ML projectsBSc/MSc/PhD in computer science/machine learningExperience working with tabular data / predictive analyticsExperience working with “classical machine learning and deep learning” (pre-LLM)Experience working with synthetic data generation, Structured Causal Models, or physical / systems-based simulatorsBenefitsCompetitive compensation with salary and equityComprehensive health coverage, including medical, dental, vision, and 401KPaid parental leave for all new parents, inclusive of adoptive and surrogate journeysRelocation support for employees moving to join the team in one of our office locationsA mission-driven, low-ego culture that values diversity of thought, ownership, and bias toward action

Requirements

Experience with: Identifying good data sources to train and evaluate ML models, including real-world and realistic synthetic data sources Bringing data from structured and unstructured sources, as well as simulators and causal models, into formats accessible by ML models Strong fundamentals of software engineering Strong knowledge of: Python Python data processing stack (numpy, pandas, …) Familiarity with: distributed processing (e.g. Ray, Dask Spark, Beam) data storage solutions Basic ML knowledge Nice to have Contributions to open source ML projects BSc/MSc/PhD in computer science/machine learning Experience working with tabular data / predictive analytics Experience working with “classical machine learning and deep learning” (pre-LLM) Experience working with synthetic data generation, Structured Causal Models, or physical / systems-based simulators

Benefits & conditions

Competitive compensation with salary and equity Comprehensive health coverage, including medical, dental, vision, and 401K Paid parental leave for all new parents, inclusive of adoptive and surrogate journeys Relocation support for employees moving to join the team in one of our office locations A mission-driven, low-ego culture that values diversity of thought, ownership, and bias toward action

About the company

Fundamental is an AI company pioneering the future of enterprise decision-making. Founded by DeepMind alumni, Fundamental has developed NEXUS - the world’s most powerful Large Tabular Model (LTM) - purpose-built for the structured records that actually drive enterprise decisions. Backed by world class investors and trusted by Fortune 100 companies, Fundamental unlocks trillions of dollars of value by giving businesses the Power to Predict. At Fundamental, you’ll work on unprecedented technical challenges in foundation model development and build technology that transforms how the world’s largest companies make decisions. This is your opportunity to be part of a category-defining company from the ground-up. Join the team defining the future of enterprise AI.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.buscojobs.com.es
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:23 min

Exploring specialized career paths within the data science ecosystem

Julian Joseph ¡ LIVE

2:34 min

Maximizing execution memory effectively via python numpy broadcasting

Jodie Burchell ¡ LIVE

2:19 min

Scaling performance across multiple GPUs using specialized frameworks

Paul Graham Paul Graham ¡ World Congress 2025

2:03 min

Accelerating pandas dataframes using cudf module plugins

Ankit Patel Ankit Patel ¡ World Congress 2024

1:50 min

Introduction to the speaker and data science background

Bas Geerdink ¡ LIVE

1:25 min

Replacing NumPy with cuPy for straightforward GPU acceleration

Paul Graham Paul Graham ¡ World Congress 2025

Videos

See all

Related articles

See all