> Markdown version of [/jobs/ext/2074163-senior-data-scientist](https://www.wearedevelopers.com/jobs/ext/2074163-senior-data-scientist). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Data Scientist - **Company:** Cloudera, Inc. - **Location:** Palo Alto, United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Automated Storage and Retrieval Systems, Code Review, Data Cleansing, Data Infrastructure, Relational Databases, Github, Iterative and Incremental Development, Python (Programming Language), Machine Learning, NumPy, Cloudera, SQL Databases, Systems Integration, Workflow Management Systems, Enterprise Software Applications, Feature Engineering, Large Language Models, Prompt Engineering, Generative AI, Jupyter, Fastapi, Pandas, AI Platforms, Core Data, Scikit Learn, Information Technology, Api Design, Streamlit Framework, Software Version Control, Unsupervised Learning - **Published:** August 15, 2026 - **Apply:** https://arc.dev/remote-jobs/j/redirect/ktck335drv ## About the Role To succeed in this role, you will demonstrate technical depth, intellectual curiosity, and a strong builder mindset: * Generative AI & LLM Engineering: Hands-on experience working with large language models (LLMs) and modern AI tooling. This includes prompt design, structured output generation, retrieval-augmented generation (RAG), evaluation strategies, and workflow automation. Ability to translate GenAI capabilities into reliable, enterprise-ready solutions that integrate with existing systems and data sources. * AI Application Development Experience: rapidly prototyping and iterating on internal applications, copilots, or AI-enabled workflow tools. Comfortable evolving prototypes into maintainable, production-grade solutions. Familiarity with modern development frameworks (e.g., Streamlit, Gradio, FastAPI, or similar) is beneficial. * Platform-Oriented Thinking: Demonstrated ability to design reusable components such as shared prompt libraries, retrieval pipelines, evaluation frameworks, and standardized integration patterns that enable scalable AI adoption. * Data Science & Machine Learning Expertise: Proficiency in Python (or R) for data preparation, feature engineering, statistical modeling, and machine learning. Experience with core data science libraries (e.g., Pandas, NumPy, scikit-learn) and a solid understanding of supervised and unsupervised learning methods. * Strong Mathematical and Statistical Foundation: Deep understanding of probability, statistical inference, experimentation, and quantitative reasoning to ensure model robustness and reliability. * SQL & Data Fluency: Strong understanding of relational databases and the ability to quickly learn new schemas and data environments. Comfortable writing efficient, production-grade SQL to support modeling, experimentation, and AI-enabled applications. * Exceptional Communication Skills: Ability to translate complex business challenges into technical solutions and clearly communicate findings, trade-offs, and recommendations to both technical and non-technical stakeholders. * Collaborative Development Experience: Experience working in collaborative environments such as Cloudera Data Science Workbench, Jupyter, Zeppelin, or similar platforms. * GitHub Proficiency: Experience using version control to support collaboration, code review, documentation, and long-term maintainability., * Hands-on experience building applications or workflows powered by large language models (LLMs). * Evidence of a builder mindset through shipped AI tools, internal platforms, or automation solutions. * Demonstrated experience applying machine learning techniques in production or enterprise environments. * 5+ years of relevant experience in Data Science, Machine Learning, or AI-focused roles. * Strong curiosity for emerging AI technologies and the ability to evaluate and adopt them responsibly. * Academic background in a quantitative discipline such as Statistics, Mathematics, Computer Science, Engineering, Economics, or a related field. You may also have: (Preferred Qualifications) * Experience with vector databases, embedding models, or semantic retrieval systems. * Experience designing internal AI platforms or shared enablement frameworks. * Familiarity with API-driven architectures and integrating AI capabilities into enterprise systems. * Exposure to responsible AI practices, governance frameworks, or model lifecycle management. ## Description You will apply rigorous analytical thinking and modern AI capabilities to design, build, and scale high-impact solutions. * Design, develop, and deploy GenAI-powered internal applications, copilots, and workflow accelerators. * Build reusable AI components, including retrieval pipelines, structured prompting patterns, orchestration workflows, and evaluation harnesses. * Design retrieval strategies that connect LLMs to trusted internal knowledge sources, ensuring grounded and reliable outputs. * Develop and maintain statistical and machine learning models to support automation, optimization, forecasting, and classification use cases. * Implement evaluation and validation frameworks to measure quality, accuracy, and consistency of AI-driven systems. * Partner cross-functionally to identify high-value opportunities for AI enablement across the organization. * Create reusable datasets, feature pipelines, and experimentation frameworks to support iterative development. * Uphold high standards for quality, reliability, and responsible AI practices. * Contribute to peer review processes to ensure technical rigor and maintainability. * Document methodologies, assumptions, and implementation details to ensure transparency and reproducibility. ## Related Videos - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [Vectorize all the things! Using linear algebra and NumPy to make your Python code lightning fast.](https://www.wearedevelopers.com/videos/562-vectorize-all-the-things-using-linear-algebra-and-numpy-to-make-your-python-code-lightning-fast) - [Advanced Typing in TypeScript](https://www.wearedevelopers.com/videos/496-advanced-typing-in-typescript) - [Coffee with Developers - Maria Apazoglou](https://www.wearedevelopers.com/videos/1209-coffee-with-developers-maria-apazoglou) - [Bringing AI Model Testing and Prompt Management to Your Codebase with GitHub Models](https://www.wearedevelopers.com/videos/1536-bringing-ai-model-testing-and-prompt-management-to-your-codebase-with-github-models) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift)