> Markdown version of [/jobs/ext/3022356-machine-learning-engineer-ai-platform-agentic-apps](https://www.wearedevelopers.com/jobs/ext/3022356-machine-learning-engineer-ai-platform-agentic-apps). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Machine Learning Engineer, AI Platform & Agentic Apps - **Company:** Robinhood - **Location:** Menlo Park, CA, United States - **Experience:** Expert - **Salary:** $255,000.0 - $300,000.0 - **Contract:** Permanent contract - **Skills:** Code Review, Python (Programming Language), Machine Learning, Management of Software Versions, Large Language Models, Information Technology - **Published:** September 21, 2026 - **Apply:** https://startup.jobs/senior-machine-learning-engineer-ai-platform-agentic-apps-robinhoodapp-10134824 ## About the Role * 10+ years of experience as a Machine Learning Engineer or ML-focused software engineer, with strong Python and distributed-systems fundamentals and a track record of shipping LLM-powered systems to production at scale. A Master's degree in Computer Science or a related technical field, or equivalent professional experience. * Hands-on experience building agentic systems end to end - tool use, orchestration, context management, multi-step planning - on top of frontier models, in production. * Deep expertise evaluating agents: you've built trajectory-level evals, tool-call scoring, and simulation environments, and you can articulate why final-answer accuracy is insufficient for systems that act. * Demonstrated expertise designing action-level guardrails - permission and tool-scoping models, approval gates, blast-radius controls, and sandboxing - for agents operating in systems where mistakes have consequences. * Rigor in evaluation methodology: golden datasets, rubric and LLM-as-judge grading and their failure modes, statistical significance with small N, offline-to-online metric correlation, and eval data versioning and contamination control. * Proven ability to build platforms, not just models: you've shipped eval, safety, or agent tooling that other engineering teams adopted, and you have the judgment to know when to build versus buy. ## Description * Design and build the core of Robinhood's agent harness - orchestration, tool integrations, context and memory management - so one platform can safely power both high-trust internal agents and tightly scoped customer-facing ones. * Ship agentic applications end to end on that harness, from an ambiguous problem to a production agent that takes real action on behalf of employees or customers, and feed what you learn back into the platform. * Build trajectory-level evaluation systems that score how an agent got to an answer, not just the answer - tool-call correctness, planning and recovery, multi-step task completion - backed by simulation environments and synthetic task generation. * Architect action guardrails as platform primitives: least-privilege tool scoping, permission models, human-approval gates for high-risk or irreversible actions, step and budget limits, sandboxing, and rollback. * Make evals and guardrails products other teams adopt - SDKs, CI regression gates on prompt, model, and tool changes, continuous red-teaming, and production tracing that closes the loop from real traffic back into eval sets and guardrail models. * Set the technical bar through architecture reviews, code reviews, and mentorship, and be the person who can make - and defend with data - the "don't ship" call. ## Related Videos - [Are Code Reviews Worth It? Insights from 16 Years of Review Data](https://www.wearedevelopers.com/videos/1135-are-code-reviews-worth-it-insights-from-16-years-of-review-data) - [Introduction to Azure Machine Learning](https://www.wearedevelopers.com/videos/368-introduction-to-azure-machine-learning) - [Creating Industry ready solutions with LLM Models](https://www.wearedevelopers.com/videos/899-creating-industry-ready-solutions-with-llm-models) - [Agentic employees in world's most downloaded FinTech app](https://www.wearedevelopers.com/videos/100123-agentic-employees-in-world-s-most-downloaded-fintech-app) - [Teaching an LLM to review code … like a Senior Engineer!](https://www.wearedevelopers.com/videos/100324-teaching-an-llm-to-review-code-like-a-senior-engineer) - [Bringing the power of AI to your application.](https://www.wearedevelopers.com/videos/1010-bringing-the-power-of-ai-to-your-application) ## Related Articles - [What is Agentic Programming and Why Should Developers Care?](https://www.wearedevelopers.com/magazine/625-what-is-agentic-programming-and-why-should-developers-care) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Never delegate the understanding](https://www.wearedevelopers.com/magazine/749-never-delegate-the-understanding) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it)