Machine Learning Engineer, AI Platform & Agentic Apps
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
Job description
- Design and build the core of Robinhoodâs agent harness - orchestration, tool integrations, context and memory management - so one platform can safely power both high-trust internal agents and tightly scoped customer-facing ones.
- Ship agentic applications end to end on that harness, from an ambiguous problem to a production agent that takes real action on behalf of employees or customers, and feed what you learn back into the platform.
- Build trajectory-level evaluation systems that score how an agent got to an answer, not just the answer - tool-call correctness, planning and recovery, multi-step task completion - backed by simulation environments and synthetic task generation.
- Architect action guardrails as platform primitives: least-privilege tool scoping, permission models, human-approval gates for high-risk or irreversible actions, step and budget limits, sandboxing, and rollback.
- Make evals and guardrails products other teams adopt - SDKs, CI regression gates on prompt, model, and tool changes, continuous red-teaming, and production tracing that closes the loop from real traffic back into eval sets and guardrail models.
- Set the technical bar through architecture reviews, code reviews, and mentorship, and be the person who can make - and defend with data - the âdonât shipâ call.
Requirements
- 10+ years of experience as a Machine Learning Engineer or ML-focused software engineer, with strong Python and distributed-systems fundamentals and a track record of shipping LLM-powered systems to production at scale. A Masterâs degree in Computer Science or a related technical field, or equivalent professional experience.
- Hands-on experience building agentic systems end to end - tool use, orchestration, context management, multi-step planning - on top of frontier models, in production.
- Deep expertise evaluating agents: youâve built trajectory-level evals, tool-call scoring, and simulation environments, and you can articulate why final-answer accuracy is insufficient for systems that act.
- Demonstrated expertise designing action-level guardrails - permission and tool-scoping models, approval gates, blast-radius controls, and sandboxing - for agents operating in systems where mistakes have consequences.
- Rigor in evaluation methodology: golden datasets, rubric and LLM-as-judge grading and their failure modes, statistical significance with small N, offline-to-online metric correlation, and eval data versioning and contamination control.
- Proven ability to build platforms, not just models: youâve shipped eval, safety, or agent tooling that other engineering teams adopted, and you have the judgment to know when to build versus buy.
Benefits & conditions
- Challenging, high-impact work to grow your career
- Performance driven compensation with multipliers for outsized impact, bonus programs, equity ownership, and 401(k) matching
- Top Tier benefits to fuel your work, including 100% paid health insurance for employees with 90% coverage for dependents
- Access to the Robinhood Employee Fund that gives eligible US employees the opportunity to invest in a private employee fund that provides exposure to Robinhood Ventures funds.
- Access to the best AI tools on the market and continuous AI skill-building for every employee, technical or not.
- Lifestyle wallet - a highly flexible benefits spending account for wellness, learning, and more
- Employer-paid life & disability insurance, fertility benefits, and mental health benefits
- Time off to recharge including company holidays, paid time off, sick time, parental leave, and more!
- Exceptional office experience with catered meals, events, and comfortable workspaces.
In addition to the base pay range listed below, this role is also eligible for bonus opportunities + equity + benefits.
Base pay for the successful applicant will depend on a variety of job-related factors, which may include education, training, experience, location, business needs, or market demands. The expected base pay range for this role is based on the location where the work will be performed and is aligned to one of 3 compensation zones. For other locations not listed, compensation can be discussed with your recruiter during the interview process.
Base Pay Range: Zone 1 (Menlo Park, CA; New York, NY; Bellevue, WA; Washington, DC) $255,000-$300,000 USD Zone 2 (Denver, CO; Westlake, TX; Chicago, IL) $225,000-$264,000 USD Zone 3 (Lake Mary, FL; Clearwater, FL; Gainesville, FL) $199,000-$234,000 USD
About the company
We are building an elite team, applying frontier technologies to the worldâs biggest financial problems. Weâre looking for bold thinkers. Sharp problem-solvers. Builders who are wired to make an impact. Robinhood isnât a place for complacency, itâs where ambitious people do the best work of their careers. Weâre a high-performing, fast-moving team with ethics at the center of everything we do. Expectations are high, and so are the rewards.
The AI Platform & Agentic Apps team builds the agent platform behind every AI agent at Robinhood. Today it gives a growing number of engineers and employees an AI teammate that ships code, queries data, and runs operational workflows on their behalf. Weâre building toward the same platform powering the agents millions of customers interact with directly, in real time. These agents donât just answer questions - theyâre designed to take real action across carefully curated meta harnesses. This is agentic AI at real scale, in a regulated financial environment, and it will change how Robinhood works!
As a Staff Machine Learning Engineer on the AI Platform & Agentic Apps team, you will design and build the harness that every agent at Robinhood runs on. A critical part of the role is making those agents trustworthy at scale: trajectory-level evals that measure how an agent reasons and acts, and action guardrails - permission models, approval gates, and sandboxing - built as platform primitives that other teams adopt. Youâll be a technical anchor on a growing, high-caliber team, collaborating with product, infrastructure, and fellow ML engineers to take ambitious ideas from zero to one and into production. Youâll help define the teamâs technical direction, mentor engineers, and shape how Robinhood decides an agent is ready to ship. This role offers a rare combination of technical depth, platform-scale impact, and the satisfaction of building systems that genuinely donât exist anywhere else.
This role is based in our Menlo Park, CA office, with in-person attendance expected at least 3 days per week.
At Robinhood, we believe in the power of in-person work to accelerate progress, spark innovation, and strengthen community. Our office experience is intentional, energizing, and designed to fully support high-performing teams.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role â technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
How to Become an AI Engineer
What Are Large Language Models?
MLOps And AI Driven Development
Never delegate the understanding