> Markdown version of [/jobs/ext/2687411-applied-scientist-ai-platform](https://www.wearedevelopers.com/jobs/ext/2687411-applied-scientist-ai-platform). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Applied Scientist - AI Platform - **Company:** Datadog - **Location:** Paris, France - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Training Data, Artificial Intelligence, Data Analysis, Distributed Systems, Monitoring of Systems, Python (Programming Language), Machine Learning, Datadog, Large Language Models, Machine Learning Operations, Gensim - **Published:** September 3, 2026 - **Apply:** https://careers.datadoghq.com/detail/8172855/?gh_jid=8172855 ## About the Role * You have a PhD, MS or equivalent research experience in a scientific field, with strong applied mathematics grounding. * 6+ years of relevant applied science or ML engineering experience, including setting technical direction for others. * You have hands-on experience with LLM and agent post-training data: how it is created, managed, and how training-data quality is controlled. This is the requirement that matters most. * You have real domain expertise in LLMs and agentic applications — not classical ML fine-tuning. Fine-tuning classifiers or traditional models is a different problem from the one this team is solving. * You have evaluated agents or LLM applications, and can define what 'good' means before you measure it. * You are a strong programmer and production software engineer. Python at minimum, plus the ability to ship scalable production systems and work with distributed systems. * You collaborate well across engineering and science teams, and you're comfortable being the domain expert who decides what comes next. * You thrive in ambiguity and can make sound technical calls when the path isn't yet defined. Bonus Points: * Hands-on LLM fine-tuning, post-training or model training experience. * Background in statistics, experiment design and data analysis. * Experience deploying production-level ML infrastructure. * Observability or monitoring systems background. * Architecture-level understanding of LLMs. ## Description * Own the applied science direction for GenSim: set the methodology and the forward-looking technical calls on how simulated environments and post-training data should be built, on a team where that decision-making does not exist yet. * Define, measure and raise the quality of post-training data — basic correctness, representativeness against the real distribution of customer systems and production telemetry, and difficulty — and make those measures something the team can act on release over release. * Close the realism gap. Simulated environments today are too clean and the injected problems are not yet hard enough; you'll drive the research and the engineering that make them look like real, imperfect production systems. * Build scalable, production-grade systems rather than research scripts. The output is not just a dataset — it is a system of synthetic environments that must be reliable and invokable inside a training loop. * Determine how this data is best applied, in LLM post-training and in evaluation, and own the agent and LLM application evaluation approaches for these environments. * Work cross-functionally with the engineers and applied scientists on adjacent teams — Bits AI SRE, the model training effort, and the wider evaluation and experimentation pillar — so that what you learn moves freely in both directions. ## Related Videos - [LLMs in the wild: Building an AI agent that survives production](https://www.wearedevelopers.com/videos/100319-llms-in-the-wild-building-an-ai-agent-that-survives-production) - [Fireside Chat: Deep Learning, Deep Impact: Harnessing AI for Language Innovation](https://www.wearedevelopers.com/videos/612-fireside-chat-deep-learning-deep-impact-harnessing-ai-for-language-innovation) - [Debugging in the Dark](https://www.wearedevelopers.com/videos/1658-debugging-in-the-dark) - [A beginner’s guide to modern natural language processing](https://www.wearedevelopers.com/videos/858-a-beginner-s-guide-to-modern-natural-language-processing) - [Coffee with Developers - Maria Apazoglou](https://www.wearedevelopers.com/videos/1209-coffee-with-developers-maria-apazoglou) - [Exploring 5 Key Applications of AI Abundance with Blockchain Assurance](https://www.wearedevelopers.com/videos/971-exploring-5-key-applications-of-ai-abundance-with-blockchain-assurance) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development)