> Markdown version of [/jobs/ext/584821-machine-learning-engineer-llm-post-training](https://www.wearedevelopers.com/jobs/ext/584821-machine-learning-engineer-llm-post-training). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Machine Learning Engineer, LLM Post-Training - **Company:** LGA AIRPORT RESTAURANTS, L.P. - **Location:** Mountain View, CA, United States - **Salary:** $150,000.0 - $230,000.0 - **Contract:** Permanent contract - **Skills:** Computer Clusters, Data Cleansing, Information Engineering, Software Debugging, Distributed Computing Environment, Machine Learning, Reinforcement Learning, Supervised Learning, Pytorch, Large Language Models, HuggingFace, Production Code, Data Generation - **Published:** June 12, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=40c075a07d429ffb ## About the Role Do you have experience in Supervised learning?, * Hands-on LLM post-training experience. You have personally run CPT, SFT, and RL training - with demonstrated, practical RL experience (RLHF / PPO / GRPO / DPO or similar), beyond just launching training scripts. * Strong data engineering for ML. You can independently design data-preparation plans for a given business scenario - sourcing, cleaning, filtering, labeling strategy, and synthetic/preference data generation - to meet specific product requirements. * Proven large-scale GPU training ability. You have trained LLMs on mid-to-large GPU hardware and are comfortable with distributed training and debugging at scale. * Strong PyTorch fundamentals; working familiarity with frameworks such as Hugging Face TRL/Accelerate, DeepSpeed or FSDP, and inference engines like vLLM. * Solid understanding of tokenization, attention, chat templates, and common failure modes in alignment/agent training. * A bias toward fast iteration and business impact, with strong communication skills to work across research and product teams., * Experience designing reward models or rule-based verifiers for RL. * Experience with tool-use / agentic model training (function calling, multi-step planning). * Publications or open-source contributions in LLM post-training or RL. ## Description We are looking for a hands-on Machine Learning Engineer to drive the post-training of our large language models, with a strong emphasis on reinforcement learning (RL). You will own the full post-training stack - continuous pre-training (CPT), supervised fine-tuning (SFT), and RL - along with the data preparation that powers it. Just as important, you will work directly with product and business teams to translate real-world use cases into concrete training objectives and ship model improvements quickly. This is a high-ownership role for someone who has actually trained models, not just read about it., * Lead post-training of our LLMs across the full pipeline: continuous pre-training, SFT, and reinforcement learning, with RL as the primary focus (e.g., RLHF, PPO, GRPO, DPO, and related methods). * Design, build, and curate the data that drives each training stage - instruction/SFT datasets, preference pairs, reward signals, on-policy rollouts, and rejection-sampled completions - and define data-preparation strategies tailored to specific business needs. * Partner closely with business and product stakeholders to understand their scenarios, rapidly convert requirements into training plans, and deliver targeted model capabilities on tight timelines. * Run large-scale training on mid-to-large GPU clusters, applying distributed-training techniques (data parallelism, FSDP, and where relevant tensor/pipeline parallelism) and tuning for throughput and stability. * Build and maintain evaluation and reward/verifier pipelines to measure model quality, prevent regressions, and ensure training-serving consistency. * Stay current with post-training research and turn promising techniques into working, production-ready code. ## Related Videos - [Adding knowledge to open-source LLMs](https://www.wearedevelopers.com/videos/1522-adding-knowledge-to-open-source-llms) - [RPA in the Public Sector](https://www.wearedevelopers.com/videos/86-rpa-in-the-public-sector) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Inside the Mind of an LLM](https://www.wearedevelopers.com/videos/1617-inside-the-mind-of-an-llm) - [Anomaly Detection - Using unsupervised Machine Learning for detecting anomalies in customer base](https://www.wearedevelopers.com/videos/6-anomaly-detection-using-unsupervised-machine-learning-for-detecting-anomalies-in-customer-base) - [Serverless deployment of (large) NLP models ](https://www.wearedevelopers.com/videos/158-serverless-deployment-of-large-nlp-models) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [The Best Large Language Models on The Market](https://www.wearedevelopers.com/magazine/319-the-best-large-language-models-on-the-market) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development)