> Markdown version of [/jobs/ext/2714701-ml-researcher](https://www.wearedevelopers.com/jobs/ext/2714701-ml-researcher). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # ML Researcher - **Company:** KREA LLC - **Location:** San Francisco, CA, United States - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Distributed Computing Environment, Open Source Technology, Reinforcement Learning, Pytorch, Large Language Models, Model Validation, Optimization Algorithms - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/ml-researcher-posttraining-krea-9836935 ## About the Role * Proven work of posttraining diffusion models for image or video generation. * Experience with large-scale model training, inference, and optimization. * Strong understanding of both LLM and diffusion post training pipelines and algorithms such as PPO, GRPO, DPO, OPD, and MOPD. * Strong proficiency in PyTorch and understanding of its inner workings. * Strong background in distributed training paradigms such as FSDP, CP, SP, USP, TP, and EP. Knowing how different parallelism strategies work together and their tradeoffs. * Good knowledge of low precision training / inference in FP8, NVFP4, and MXFP8. * Good understanding of algorithms and techniques used in fast inference engines such as vLLM and sglang as well as existing RL frameworks in LLM space such as slime, miles, tinker, and verl. * Understanding of various RL infrastructure and optimization techniques such as async RL, fast weight transfer, pipelining rollouts, managing off policy data. * Ability to monitor model regression and identify weak areas and turn them into concrete evals and reward design. * Experience training VLM models. Many of our custom reward models use VLM to provide reward signals for our models. * Keeping up with the developments in related fields such as LLM, VLM, representation learning, and robotics research. * Being comfortable working in a goal-oriented research environment. * Having good judgement around when one should explore different training strategies and when it's time to commit to a specific strategy to scale compute and data. * Comfortable working with underspecified goals. We expect every technical member to take an ambiguous research goal and break it down into concrete requirements, plans, experiment plan, and execution items. * Good research taste - bias towards simplicity and methods that scale well with compute, data, and minimal human supervision. ## Description * We work full-time and in-person at our North Beach office in San Francisco. * We believe that demonstrated interest in the creative space is key: our team includes musicians, designers, visual artists and more. * Fast iteration and execution speed. Bias towards action, agency, and independence. What you'll do * Finetune diffusion models at scale to improve image aesthetics and quality. * Implement posttraining techniques ranging from supervised finetuning, preference optimization, reinforcement learning, on-policy distillation, and various distillation / acceleration techniques. * Design comprehensive eval suites and reward designs for the reinforcement learning stage focused on image space. * Train custom VLM as reward models as part of our reward design. * Train custom LLMs for prompt expansion through finetuning and reinforcement learning. * Coordinate with data teams and partners to manage collection of preference data and model evaluation results. * Work on safety alignment of our models for open source release. * Collaborate with our AI research and engineering teams to integrate advancements into our products. ## Related Videos - [Creating Industry ready solutions with LLM Models](https://www.wearedevelopers.com/videos/899-creating-industry-ready-solutions-with-llm-models) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Introduction to Responsible AI: Balancing Value and Risk](https://www.wearedevelopers.com/videos/1972-introduction-to-responsible-ai-balancing-value-and-risk) - [Adding knowledge to open-source LLMs](https://www.wearedevelopers.com/videos/1522-adding-knowledge-to-open-source-llms) - [Bringing the power of AI to your application.](https://www.wearedevelopers.com/videos/1010-bringing-the-power-of-ai-to-your-application) - [Lies, Damned Lies and Large Language Models](https://www.wearedevelopers.com/videos/1231-lies-damned-lies-and-large-language-models) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [The Best Large Language Models on The Market](https://www.wearedevelopers.com/magazine/319-the-best-large-language-models-on-the-market) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [Everything a Developer Needs to Know About MCP with Neo4j](https://www.wearedevelopers.com/magazine/604-everything-a-developer-needs-to-know-about-mcp-with-neo4j) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development)