> Markdown version of [/jobs/ext/1417222-data-scientist-ii](https://www.wearedevelopers.com/jobs/ext/1417222-data-scientist-ii). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Scientist II - **Company:** Robert Half - **Location:** San Francisco, CA, United States - **Experience:** Expert - **Salary:** $119,000.0 - $180,000.0 - **Contract:** Permanent contract - **Skills:** A/B Testing, Artificial Intelligence, Monitoring of Systems, Python (Programming Language), Machine Learning, Chatbots, Large Language Models, Multi-Agent Systems, Prompt Engineering, Information Technology, Automation Anywhere - **Published:** July 24, 2026 - **Apply:** https://dejobs.org/x/x/65297EC2CF224163B2F5122374834363/job/ ## About the Role * Master's or Ph.D. in Computer Science, Data Science, Machine Learning, or related fields. * 5+ years of experience in applied ML, AI development, or advanced data science roles. * Hands-on experience building and deploying LLM-based agents, RAG pipelines, and multimodal Gen AI systems. * Proficiency in Python and Gen AI frameworks * Strong understanding of agent orchestration frameworks (LangGraph, LangChain, AutoGPT, or equivalent). * Demonstrated expertise in benchmarking, A/B testing, and model monitoring for AI systems. * Experience developing and maintaining prompt engineering standards and frameworks. * Ability to evaluate and optimize model performance across different LLMs and architectures. * Strong communication and collaboration skills across product, engineering, and business teams. * Proven track record of mentoring team members and driving innovation in AI workflows. * Familiarity with internal experimentation, observability, or AI governance tools is a plus. * Contributions to open-source projects, research publications, or AI community initiatives is a plus. ## Description * Design and implement Gen AI-powered agents that solve real business challenges and streamline internal workflows. * Extend and optimize the internal agent evaluation framework to ensure reliable, high-performing, and explainable agents. * Build benchmarking and monitoring systems to assess orchestration strategies and model performance in production. * Conduct experiments to evaluate the impact of model or architecture changes, balancing performance, cost, and scalability. * Build on existing prompt engineering frameworks to standardize best practices and eliminate "prompt debt." * Move beyond chatbots to create ambient, context-aware agents that boost efficiency and unlock new productivity avenues. * Collaborate with engineering teams to advance the internal agentic platform, focusing on orchestration, observability, and optimization. * Mentor and guide junior data scientists and software engineers on Gen AI techniques and frameworks. * Partner with product managers to assess technical feasibility, estimate level of effort (LoE), and prioritize development initiatives. * Bring innovative ideas to evolve agentic architecture and continuously advance internal AI capabilities. * Communicate complex AI concepts effectively to non-technical stakeholders, supporting organizational learning and adoption. * Lead roadshows and demos to advocate for Gen AI tools and drive cultural change across the enterprise. ## Related Videos - [Bringing the power of AI to your application.](https://www.wearedevelopers.com/videos/1010-bringing-the-power-of-ai-to-your-application) - [Chatbots are going to destroy infrastructures and your cloud bills](https://www.wearedevelopers.com/videos/1130-chatbots-are-going-to-destroy-infrastructures-and-your-cloud-bills) - [Introduction to Azure Machine Learning](https://www.wearedevelopers.com/videos/368-introduction-to-azure-machine-learning) - [How AI Models Get Smarter](https://www.wearedevelopers.com/videos/1374-how-ai-models-get-smarter) - [Testing AI Agents: Automated Evaluation for Chatbots & RAG Systems](https://www.wearedevelopers.com/videos/100300-testing-ai-agents-automated-evaluation-for-chatbots-rag-systems) - [What non-automotive Machine Learning projects can learn from automotive Machine Learning projects](https://www.wearedevelopers.com/videos/397-what-non-automotive-machine-learning-projects-can-learn-from-automotive-machine-learning-projects) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [What Industries Outside of AI Are Hiring The Most AI Experts?](https://www.wearedevelopers.com/magazine/98-what-industries-outside-of-ai-are-hiring-the-most-ai-experts) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers)