> Markdown version of [/jobs/ext/2826331-junior-data-scientist](https://www.wearedevelopers.com/jobs/ext/2826331-junior-data-scientist). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Junior Data Scientist - **Company:** KAHUNA LLC - **Location:** United States (Remote available) - **Experience:** Starter - **Contract:** Permanent contract - **Skills:** Training Data, Application Programming Interfaces (APIs), Amazon Web Services, Data Analysis, Databases, Data Cleansing, Data Visualization, Monitoring of Systems, Python (Programming Language), PostgreSQL, Machine Learning, NumPy, Raw Data, SQL Databases, Feature Engineering, Deep Learning, Model Validation, Git, Fastapi, Pandas, Matplotlib, Scikit Learn, Xgboost, Plotly, Data Pipelines, Docker - **Published:** September 10, 2026 - **Apply:** https://arc.dev/remote-jobs/j/redirect/pj24q9c17q ## About the Role * At least 1 year of hands-on experience in data science, machine learning, or a closely related analytical role. * Strong working knowledge of Python for data analysis. * Hands-on experience with pandas, NumPy, and scikit-learn or similar libraries. * Practical experience with the machine learning workflow, including: * data cleaning and exploratory analysis, * feature engineering, * train / validation / test design, * model selection and training, * model evaluation, diagnostics, and error analysis. * Understanding of supervised machine learning methods, particularly regression, classification, and tree-based models. * Ability to work with SQL and relational datasets. * Experience creating analytical visualisations with Matplotlib, Plotly, or similar libraries. * Ability to identify and investigate issues involving missing data, outliers, duplicates, inconsistent sources, or unexpected results. * Comfortable using Git and working with an existing codebase. * Ability to research unfamiliar problems independently and explain the reasoning behind an approach. * Professional-level written and spoken English. Nice to Have These are useful, but not requirements: * Experience with XGBoost, LightGBM, or similar gradient boosting libraries. * Experience with time-series or forecasting problems. * Familiarity with experiment tracking or model monitoring. * Experience with APIs, FastAPI, Docker, AWS, or similar production tooling. * Experience with PostgreSQL or other production databases. * Experience working with larger datasets or data pipelines. * Interest in music, streaming platforms, or the digital music ecosystem. Production engineering experience is not expected across all of these areas. Strong data science fundamentals and the ability to learn the surrounding tooling are more important. ## Description We're looking for a Junior Data Scientist to work on applied machine learning and analytics problems across music performance, forecasting, artist discovery, and underwriting support. This is primarily a tabular data and classical machine learning role, not an audio/MIR or deep learning role. You'll work with real-world datasets, prepare analytical and training data, develop and evaluate machine learning models, investigate model behaviour, and communicate results through clear analysis and visualisation. We expect candidates to already understand the core machine learning workflow. You don't need experience with every tool in our stack, the ability to investigate unfamiliar problems, use technical documentation effectively, and learn independently matters more. What You'll Do * Explore, clean, transform, and validate streaming, royalty, artist, and platform data. * Build reproducible analytical datasets and feature engineering workflows using Python. * Develop and evaluate models for forecasting, scoring, ranking, and classification problems. * Perform model diagnostics and error analysis to understand where and why performance changes. * Create data visualisations and analytical outputs for product, scouting, underwriting, and business use cases. * Write SQL queries to retrieve, validate, and investigate data. * Work with existing data pipelines and contribute to their reliability and maintainability. * Document datasets, assumptions, experiments, model behaviour, and analytical findings. * Work with product and engineering when analytical outputs need to be integrated into internal or artist-facing products., * See the full lifecycle from raw data and experimentation to production use and monitoring. * Work on applied problems at the intersection of data science, music, and finance. * Full remote working option.