> Markdown version of [/jobs/ext/1301272-data-scientist](https://www.wearedevelopers.com/jobs/ext/1301272-data-scientist). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Scientist - **Company:** Kroll Inc - **Location:** United States (Remote available) - **Experience:** Starter - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Data Analysis, Microsoft Azure, Cloud Computing, Code Review, Data Cleansing, Python (Programming Language), Machine Learning, NumPy, Tensorflow, Azure Machine Learning, Unstructured Data, Feature Engineering, Pytorch, Retrieval-Augmented Generation, Large Language Models, Prompt Engineering, Apache Spark, Deep Learning, Model Validation, Pandas, Data Lakes, Pyspark, Scikit Learn, Information Technology, HuggingFace, Machine Learning Operations, Software Version Control, Data Pipelines, Serverless Computing, Unsupervised Learning, Databricks - **Published:** July 16, 2026 - **Apply:** https://arc.dev/remote-jobs/j/redirect/p51pyerprr ## About the Role * Bachelor's or Master's degree in computer science, statistics, mathematics, data science, or a related quantitative field * 1-3 years of practical data science or ML experience (internships, co-ops, research, and strong project work all count) * Proficiency in Python and familiarity with core ML libraries (scikit-learn, pandas, NumPy) * Solid understanding of foundational ML concepts: supervised and unsupervised learning, model evaluation, cross-validation, and feature engineering * Exposure to at least one deep learning or NLP framework (PyTorch, TensorFlow, or Hugging Face Transformers) and familiarity with LLM concepts such as prompt engineering, embeddings, or retrieval-augmented generation * Comfort working with structured and unstructured data, including text and document-based sources * Basic understanding of the ML lifecycle - from data preparation and experimentation through evaluation and handoff * Clear, organised communication skills - able to document work and explain methods to peers and stakeholders * Curiosity, rigour, and a strong drive to learn in a fast-paced, collaborative environment Preferred * Hands-on experience with Databricks, Spark/PySpark, or cloud ML platforms (Azure AI Foundry, AWS SageMaker, or GCP Vertex AI) * Hands-on experience with LLM/GenAI and agentic workflows - prompt engineering, RAG, embeddings, vector databases, or building with frameworks such as LangChain, LlamaIndex, or Semantic Kernel * Familiarity with MLflow or other experiment tracking and model versioning tools * Experience in financial services, risk, compliance, or a regulated industry * Knowledge of responsible AI principles, including fairness, transparency, and data privacy ## Description * Build, train, and evaluate machine learning models across traditional ML, NLP, and LLM/GenAI use cases, under the guidance of senior team members * Contribute to data pipelines and feature engineering workflows in Databricks using PySpark and Delta Lake * Support model deployment and monitoring on Azure - including Azure AI Foundry, Azure OpenAI, and Azure Functions - and help maintain production model health * Contribute to LLM and generative AI workflows - including prompt engineering, RAG pipelines, and agentic applications built on frameworks such as LangChain or LlamaIndex - under the guidance of senior team members * Conduct exploratory data analysis and communicate findings clearly through code, documentation, and team presentations * Write clean, well-tested, and reproducible Python code; contribute to shared codebases and track experiments via MLflow * Partner with senior data scientists and cross-functional stakeholders to understand business problems and translate them into analytical approaches * Participate actively in code reviews, team rituals, and knowledge-sharing sessions * Develop your skills proactively - engage with new tools, research, and techniques relevant to the team's work ## Related Videos - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [Vectorize all the things! Using linear algebra and NumPy to make your Python code lightning fast.](https://www.wearedevelopers.com/videos/562-vectorize-all-the-things-using-linear-algebra-and-numpy-to-make-your-python-code-lightning-fast) - [Advanced Typing in TypeScript](https://www.wearedevelopers.com/videos/496-advanced-typing-in-typescript) - [Empowering Retail Through Applied Machine Learning](https://www.wearedevelopers.com/videos/976-empowering-retail-through-applied-machine-learning) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers)