Data Scientist

KANINI LLC
United States
2 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
2 years minimum
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Airflow Amazon Web Services Amazon S3 Data Analysis ARM Architecture BigQuery Cloud Computing Cluster Analysis Computer Programming Databases Continuous Integration
+44 more
Data Governance Data Infrastructure Data Transformation Data Profiling Distributed Computing Environment Statistical Hypothesis Testing Python (Programming Language) Machine Learning Natural Language Processing NumPy Query Optimization Tensorflow Azure Machine Learning SQL Databases Reinforcement Learning Feature Engineering Pytorch Prophet Flask (Web Framework) Large Language Models Snowflake Prompt Engineering Apache Spark Deep Learning Generative AI Fastapi Pandas Matplotlib Pyspark Core Data Scikit Learn Kubernetes Information Technology Statistics Packages HuggingFace Data Analytics Xgboost Plotly Machine Learning Operations Restful APIs Data Pipelines Recurrent Neural Networks Docker Unsupervised Learning

Job description

We are hiring a pioneering, fully autonomous Senior Data Scientist to own the complete data science lifecycle for our flagship prediction model project. This role is mission-critical - the candidate will be the single point of expertise responsible for sourcing, analyzing, and engineering all data that powers our predictive models. Operating independently with minimal supervision, this individual must combine deep AI/ML mastery, hands-on engineering skills, and sharp business acumen to deliver measurable, production-grade outcomes., Data Analysis & Pipeline Ownership

  • Lead end-to-end analysis of large, complex, multi-source datasets to surface patterns driving model inputs

  • Identify, collect, clean, validate, and transform all data required for prediction model consumption

  • Design and maintain scalable, production-grade data pipelines (training, validation, inference)

  • Perform deep EDA, data profiling, and quality audits to ensure model-ready data standards

Predictive Modeling & AI/ML

  • Architect, train, evaluate, and iterate ML models - supervised, unsupervised, and reinforcement learning

  • Own feature engineering: selection, extraction, transformation, and dimensionality reduction

  • Apply advanced techniques: deep learning, NLP, time-series forecasting, ensemble methods

  • Benchmark, A/B test, and monitor models in production; drive continuous performance improvement

  • Deploy models via REST APIs (FastAPI/Flask); ensure reproducibility and scalability

Independent Ownership & Leadership

  • Self-direct from problem definition through solution delivery with zero hand-holding

  • Translate ambiguous business problems into precise, executable data science problem statements

  • Communicate model results and data insights clearly to technical and non-technical stakeholders

  • Document all experiments, methodologies, and outcomes - audit-ready and reproducible

  • Champion best practices across the data science lifecycle; mentor junior team members

Requirements

  • B.S./M.S./Ph.D. in Computer Science, Statistics, Mathematics, or equivalent quantitative field (Master’s/Ph.D. strongly preferred)

  • 5+ years of hands-on data science experience with at least 2 years delivering production-grade ML models

  • Proven ability to own and deliver end-to-end data science projects independently

  • Portfolio demonstrating innovation in predictive modeling and measurable business impact

  • Experience in a fast-paced, data-driven, decision-model environment, Core Data Science & Mathematics

  • Statistics (Bayesian inference, hypothesis testing, regression, distributions)

  • Linear algebra, calculus, and probability applied to ML model design

  • Supervised & unsupervised learning, anomaly detection, clustering

  • Time-series analysis & forecasting: ARIMA, Prophet, LSTM

Programming & Development

  • Python (Expert): NumPy, Pandas, Scikit-learn, Statsmodels, Matplotlib, Plotly

  • SQL (Advanced): window functions, CTEs, query optimization

  • Git / GitHub; CI/CD for ML; MLOps with MLflow or Kubeflow

  • Docker & Kubernetes for model containerization and serving

AI / ML Frameworks (Must-Have)

  • TensorFlow and/or PyTorch - deep learning architectures

  • XGBoost, LightGBM, CatBoost - gradient boosting & ensemble methods

  • Hugging Face Transformers - NLP, LLMs, and fine-tuning

  • SHAP, LIME - model explainability and interpretability

  • LLMs / Generative AI / Prompt Engineering - strong advantage

Cloud & Data Infrastructure

  • AWS (SageMaker, S3, Glue), GCP (Vertex AI, BigQuery), or Azure ML

  • Apache Spark / PySpark - distributed data processing

  • Airflow / Prefect - pipeline orchestration

Snowflake (Good to Have)

  • Snowflake Data Cloud: querying, Snowpark for Python ML pipelines

  • Snowflake Cortex AI / ML Functions for in-database ML

  • dbt for data transformation; data governance within Snowflake

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on arc.dev

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

2:03 min

Accelerating pandas dataframes using cudf module plugins

Ankit Patel Ankit Patel · WWC 2024

2:28 min

Identifying root causes through global and local SHAP plots

Bernhard Bernhard +1 · WWC 2025

2:34 min

Maximizing execution memory effectively via python numpy broadcasting

Jodie Burchell · LIVE

9:52 min

Live Matplotlib rendering and data plotting within Deepnote environments

Radovan Kavický · LIVE

3:33 min

Refactoring data science workflows using Rapids QDF and Pandas

Paul Graham Paul Graham · LIVE

Videos

See all

Related articles

See all