> Markdown version of [/jobs/ext/2003342-data-scientist](https://www.wearedevelopers.com/jobs/ext/2003342-data-scientist). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Scientist - **Company:** KANINI LLC - **Location:** United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Airflow, Amazon Web Services, Amazon S3, Data Analysis, ARM Architecture, BigQuery, Cloud Computing, Cluster Analysis, Computer Programming, Databases, Continuous Integration, Data Governance, Data Infrastructure, Data Transformation, Data Profiling, Distributed Computing Environment, Statistical Hypothesis Testing, Python (Programming Language), Machine Learning, Natural Language Processing, NumPy, Query Optimization, Tensorflow, Azure Machine Learning, SQL Databases, Reinforcement Learning, Feature Engineering, Pytorch, Prophet, Flask (Web Framework), Large Language Models, Snowflake, Prompt Engineering, Apache Spark, Deep Learning, Generative AI, Fastapi, Pandas, Matplotlib, Pyspark, Core Data, Scikit Learn, Kubernetes, Information Technology, Statistics Packages, HuggingFace, Data Analytics, Xgboost, Plotly, Machine Learning Operations, Restful APIs, Data Pipelines, Recurrent Neural Networks, Docker, Unsupervised Learning - **Published:** August 9, 2026 - **Apply:** https://arc.dev/remote-jobs/j/redirect/paepzpc06l ## About the Role * B.S./M.S./Ph.D. in Computer Science, Statistics, Mathematics, or equivalent quantitative field (Master's/Ph.D. strongly preferred) * 5+ years of hands-on data science experience with at least 2 years delivering production-grade ML models * Proven ability to own and deliver end-to-end data science projects independently * Portfolio demonstrating innovation in predictive modeling and measurable business impact * Experience in a fast-paced, data-driven, decision-model environment, Core Data Science & Mathematics * Statistics (Bayesian inference, hypothesis testing, regression, distributions) * Linear algebra, calculus, and probability applied to ML model design * Supervised & unsupervised learning, anomaly detection, clustering * Time-series analysis & forecasting: ARIMA, Prophet, LSTM Programming & Development * Python (Expert): NumPy, Pandas, Scikit-learn, Statsmodels, Matplotlib, Plotly * SQL (Advanced): window functions, CTEs, query optimization * Git / GitHub; CI/CD for ML; MLOps with MLflow or Kubeflow * Docker & Kubernetes for model containerization and serving AI / ML Frameworks (Must-Have) * TensorFlow and/or PyTorch - deep learning architectures * XGBoost, LightGBM, CatBoost - gradient boosting & ensemble methods * Hugging Face Transformers - NLP, LLMs, and fine-tuning * SHAP, LIME - model explainability and interpretability * LLMs / Generative AI / Prompt Engineering - strong advantage Cloud & Data Infrastructure * AWS (SageMaker, S3, Glue), GCP (Vertex AI, BigQuery), or Azure ML * Apache Spark / PySpark - distributed data processing * Airflow / Prefect - pipeline orchestration Snowflake (Good to Have) * Snowflake Data Cloud: querying, Snowpark for Python ML pipelines * Snowflake Cortex AI / ML Functions for in-database ML * dbt for data transformation; data governance within Snowflake ## Description We are hiring a pioneering, fully autonomous Senior Data Scientist to own the complete data science lifecycle for our flagship prediction model project. This role is mission-critical - the candidate will be the single point of expertise responsible for sourcing, analyzing, and engineering all data that powers our predictive models. Operating independently with minimal supervision, this individual must combine deep AI/ML mastery, hands-on engineering skills, and sharp business acumen to deliver measurable, production-grade outcomes., Data Analysis & Pipeline Ownership * Lead end-to-end analysis of large, complex, multi-source datasets to surface patterns driving model inputs * Identify, collect, clean, validate, and transform all data required for prediction model consumption * Design and maintain scalable, production-grade data pipelines (training, validation, inference) * Perform deep EDA, data profiling, and quality audits to ensure model-ready data standards Predictive Modeling & AI/ML * Architect, train, evaluate, and iterate ML models - supervised, unsupervised, and reinforcement learning * Own feature engineering: selection, extraction, transformation, and dimensionality reduction * Apply advanced techniques: deep learning, NLP, time-series forecasting, ensemble methods * Benchmark, A/B test, and monitor models in production; drive continuous performance improvement * Deploy models via REST APIs (FastAPI/Flask); ensure reproducibility and scalability Independent Ownership & Leadership * Self-direct from problem definition through solution delivery with zero hand-holding * Translate ambiguous business problems into precise, executable data science problem statements * Communicate model results and data insights clearly to technical and non-technical stakeholders * Document all experiments, methodologies, and outcomes - audit-ready and reproducible * Champion best practices across the data science lifecycle; mentor junior team members ## Related Videos - [Python Data Visualization @ Deepnote (w/ PyViz overview)](https://www.wearedevelopers.com/videos/113-python-data-visualization-deepnote-w-pyviz-overview) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [Advanced Typing in TypeScript](https://www.wearedevelopers.com/videos/496-advanced-typing-in-typescript) - [Vectorize all the things! Using linear algebra and NumPy to make your Python code lightning fast.](https://www.wearedevelopers.com/videos/562-vectorize-all-the-things-using-linear-algebra-and-numpy-to-make-your-python-code-lightning-fast) - [Getting to Know Your Legacy (System) with AI-Driven Software Archeology](https://www.wearedevelopers.com/videos/1437-getting-to-know-your-legacy-system-with-ai-driven-software-archeology) - [How to implement convenient Python bindings to C++](https://www.wearedevelopers.com/videos/618-how-to-implement-convenient-python-bindings-to-c) ## Related Articles - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development)