> Markdown version of [/jobs/ext/3044387-data-scientist](https://www.wearedevelopers.com/jobs/ext/3044387-data-scientist). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Scientist - **Company:** Smadex SLU - **Location:** Barcelona, Spain - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Python (Programming Language), Machine Learning, NumPy, Kaggle, Pandas, Scikit Learn, Xgboost - **Published:** September 24, 2026 - **Apply:** https://www.adzuna.es/contact-us.html ## About the Role 5+ years of experience in Data Science, Machine Learning, or applied ML engineering roles. Strong proficiency in Python, including hands-on experience with pandas, NumPy, and scikit-learn. Deep expertise in traditional machine learning models such as XGBoost, LightGBM, CatBoost, and other gradient boosting frameworks. Strong understanding of real-world tabular data challenges, including missing values, class imbalance, high cardinality features, and distribution shift. Proven experience designing experiments, building benchmarks, and drawing rigorous conclusions from noisy or imperfect data. Ability to independently drive projects from research idea to production-ready implementation. Strong analytical and problem-solving mindset with a focus on measurable impact and empirical validation. Experience working with structured prediction problems in domains such as finance, healthcare, supply chain, retail, or industrial applications is a plus. Familiarity with tabular foundation models (e.g., TabPFN, CARTE) is a strong advantage. Exposure to tools such as DuckDB, Polars, or modern in-process analytics engines is a plus. Experience participating in competitive data science environments (e.g., Kaggle, DrivenData) or contributing to ML libraries is beneficial. Ability to read, interpret, and apply machine learning research papers in practical implementations. ## Description As a Data Scientist focused on Extensions, you will help advance cutting-edge AI systems designed for enterprise decision-making at scale. You will work on complex structured data problems, improving the predictive performance of large tabular models across diverse industries and real-world use cases. This role sits at the intersection of research and production, requiring you to translate experimental ideas into robust, production-grade solutions that directly impact enterprise customers. You will collaborate closely with research, engineering, and applied AI teams to understand model behavior and enhance system capabilities. The environment is highly technical, fast-moving, and research-driven, offering the opportunity to contribute to foundational AI technology. Your work will directly influence how large organizations leverage data to make better, faster, and more accurate decisions. Accountabilities Research and develop advanced data science methods to improve predictive performance across large-scale structured enterprise datasets and diverse prediction tasks. Design, implement, and maintain production-quality Python components with a strong focus on correctness, scalability, and reusability. Analyze real-world enterprise data characteristics and design strategies to ensure robust model performance under conditions such as missing data, class imbalance, and distribution shifts. Design and run rigorous experiments, build meaningful benchmarks, and evaluate model improvements using statistically sound methodologies. Work across a broad range of structured machine learning problems including classification, regression, ranking, and forecasting. Collaborate closely with research and engineering teams to understand model behavior and translate insights into product improvements. Partner with Applied AI Engineers to validate approaches on real customer datasets and transform findings into deployable capabilities. Contribute to technical documentation, internal tooling, and best practices to improve reproducibility and knowledge sharing across teams. ## Related Videos - [Machine learning 101: Where to begin?](https://www.wearedevelopers.com/videos/1014-machine-learning-101-where-to-begin) - [Vectorize all the things! Using linear algebra and NumPy to make your Python code lightning fast.](https://www.wearedevelopers.com/videos/562-vectorize-all-the-things-using-linear-algebra-and-numpy-to-make-your-python-code-lightning-fast) - [Advanced Typing in TypeScript](https://www.wearedevelopers.com/videos/496-advanced-typing-in-typescript) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [Data Science on Software Data](https://www.wearedevelopers.com/videos/162-data-science-on-software-data) - [Python Data Visualization @ Deepnote (w/ PyViz overview)](https://www.wearedevelopers.com/videos/113-python-data-visualization-deepnote-w-pyviz-overview) ## Related Articles - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [The 13 Best Python Libraries for Developers in 2025](https://www.wearedevelopers.com/magazine/371-the-13-best-python-libraries-for-developers-in-2025) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [Data Analyst Salary in the UK](https://www.wearedevelopers.com/magazine/278-data-analyst-salary-in-the-uk) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know)