> Markdown version of [/magazine/773-i-gave-a-video-editor-more-autonomy-than-a-trading-bot-on-purpose](https://www.wearedevelopers.com/magazine/773-i-gave-a-video-editor-more-autonomy-than-a-trading-bot-on-purpose). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # I Gave a Video Editor More Autonomy Than a Trading Bot. On Purpose. **By:** Igor Khokhriakov, [Chris Heilmann](https://www.wearedevelopers.com/@chris-heilmann) **Published:** October 2, 2026 This is a guest post by Igor Khokhriakov. If you asked me which of two systems should get more AI autonomy, an algorithmic trading platform or a hobby video editor, the obvious answer would probably be the trading platform. Mine is designed the other way around. The trading system can take an idea expressed in natural language and eventually deploy a live AWS Lambda that acts on real capital. Yet architecturally, it is the less autonomous system. AWS Step Functions owns the route. The model works inside bounded stages. Tests, schemas, deployment machinery, and explicit state transitions decide what is allowed to happen next. The video system has a much smaller blast radius. If it makes a bad decision, I get a bad 60-second trailer and throw it away. So there I let the model make far more of the decisions that actually define the result: which moments matter, what the hook is, how a cut should be structured, which soundtrack fits, and what story the trailer tells. An autonomous coding agent also builds, runs, and repairs much of that pipeline under a set of contracts I defined. That led me to a rule I now use when designing agentic systems: **Grant an agent autonomy in inverse proportion to the cost of a wrong action.** The more expensive a mistake is, the less control flow I want the model to own. ## The distinction I care about “Agent” has become an overloaded word. A loop that calls an LLM and a few tools gets called an agent. So does a workflow with one model step. So does a coding system that can modify a repository for an hour. For architecture, I find a simpler question more useful: **Who owns the control flow?** In a workflow, software decides what can happen next. The model performs bounded work inside that structure. In an agent-led system, the model gets to make more of those runtime choices itself. It decides which path to pursue, which tool to use, what to retry, or how to transform an intermediate result. There can still be a deterministic envelope around an agent. In fact, there usually should be. The important variable is how much of the decision space you hand over. The two systems look roughly like this: HIGH COST OF ERROR LOW COST OF ERROR CODE OWNS ROUTE MODEL OWNS JUDGMENT │ │ ▼ ▼ [Designer] ┌──── analyze ─────┐ │ │ │ ▼ ▼ │ [Selector] choose moments │ │ │ │ ▼ ├── plan edit │ [Researcher] ├── pick music │ │ ├── write story │ ▼ └── revise ────────┘ [Deployer] │ │ ▼ tests / build / deterministic deploy actuators │ │ ▼ ▼ REAL MONEY BAD VIDEO autonomy: bounded autonomy: broad The interesting design decision is not whether an LLM appears in the system. It is where control changes hands. ## The trading pipeline: autonomy inside a cage One of my personal R&D projects is a closed-loop algorithmic trading pipeline. I can give it a strategy idea in natural language. In its current end-to-end mode, it can turn that into a specification, implementation, candidate-symbol research, parameter optimization, generated deployment code, tests, an AWS build, and finally a live Lambda deployment without a manual handoff between stages. The AWS infrastructure costs roughly **$2.50 per month**. That sounds like exactly the sort of thing people now describe as an autonomous trading agent. But the architecture is deliberately boring. AWS Step Functions owns the top-level control flow. It invokes four ECS Fargate stages: Designer, Selector, Researcher, and Deployer. A stage can perform model-driven work, but it cannot decide that the next stage no longer matters, invent a fifth production step, or jump around the state machine because the model “feels” that it has enough information. Code owns those transitions. That does not mean the model is decorative. The Designer turns intent into structured artifacts. Research involves automated experimentation. The Deployer contains an agentic coding loop that can write a strategy implementation, run its contract tests, inspect failures, modify the code, and try again. The distinction is where that autonomy stops. The coding loop has a bounded job. When it finishes, ordinary software takes over again. Contract tests must pass. Scenario tests exercise behavior. SAM builds the package. Deployment goes through the AWS machinery. Failures become explicit pipeline failures rather than conversational improvisation. If an LLM hallucinates a method name, I do not want another LLM deciding whether that hallucination is probably harmless. I want `pytest` to exit non-zero. This is the part of agent architecture that is easy to underestimate: **determinism is not the opposite of AI. Determinism is how you decide where AI is allowed to matter.** The pipeline even has human-in-the-loop gates that can be enabled per stage. In my current setup they are disabled for the end-to-end path, but the control points exist as part of the architecture rather than as an emergency feature I would have to retrofit later. Operationally, the same philosophy continues. Failures move through explicit infrastructure. Deployment state is observable. Operators have kill, pause, resume, and fleet-management controls. Components that do not need to stay alive do not stay alive. Fargate is pay-per-use, Lambda has no warm worker pool, EventBridge handles scheduled jobs, DynamoDB is on-demand, and artifacts live in S3. None of that is glamorous. That is why I trust it around money. ## The video pipeline: buy autonomy with a small blast radius The other project started from old video footage. The goal was to turn source material into vertical Shorts and roughly one-minute trailers. The input might be archival footage or a travel vlog. The interesting problem is not transcoding a video. FFmpeg solved that years ago. The interesting problem is editorial judgment. Which 20 seconds are actually worth watching? What should the first three seconds communicate? Which visual moment supports a line from the transcript? Which candidate belongs in a trailer? Which soundtrack changes the emotional reading of the same cuts? Those are fuzzy decisions, so I let the model make them. The system uses n8n as the workflow envelope and custom model-powered nodes for analysis, edit planning, visual captioning, soundtrack selection, and narration. A Highlight Analyzer proposes and ranks moments. An Edit Planner converts an approved idea into a concrete cut. A visual model describes what is actually visible. A Reel Director selects music from the available library and defines the trailer presentation. A Narrator constructs the narrative layer. The pixels themselves are not generated by the model. The model produces structured intent. Deterministic FFmpeg code executes it. There is also a second layer of autonomy. I built the project to be operated across coding-agent sessions. The repository contains contracts, acceptance criteria, runbooks, and gates. A coding agent can inspect the state of the project, run workflows, diagnose infrastructure problems, change implementation code, rerun tests, and continue toward the artifact. My role changes from typing every implementation detail to defining the system in which that work is allowed to happen. That is much closer to what I consider genuinely useful agentic engineering. And I am comfortable giving it that freedom for one simple reason: the failure is cheap. A broken trading deployment can spend money. A bad edit wastes a few minutes. That difference should radically change the architecture. ## Guard the actuator, not just the prompt Giving the video system more autonomy does not mean telling the model “do whatever you want.” One rule from this project has become more important to me than prompt engineering: **If an action must never happen, enforce that where the action happens.** The archive footage I was working with had a specific requirement: preserve the original analog character. No AI upscaling, synthetic VHS effects, sharpening, denoising, or fake glitch filters. I could have put that in the system prompt in capital letters. I did put it in the specification, but I did not stop there. The renderer itself maintains an allowlist: const ALLOWED_FILTERS = [ 'trim', 'atrim', 'setpts', 'asetpts', 'concat', 'scale', 'setsar', 'pad', 'crop', 'subtitles', 'drawtext', 'loudnorm', 'anull', 'aresample', 'format', 'fps', 'yadif', 'volume', 'afade', ]; function assertAllowedFilters(filterComplex) { const used = extractFilterNames(filterComplex); const disallowed = [...used].filter(f => !ALLOWED_FILTERS.includes(f)); if (disallowed.length) { throw new Error( `filter not on allowlist: ${disallowed.join(', ')}` ); } } Before FFmpeg runs, the generated filter graph is inspected. If something outside the safe set appears, rendering fails. The model cannot negotiate with that check. It cannot persuade it. It cannot discover a more creative interpretation of the phrase “preserve VHS authenticity.” The actuator simply cannot perform that action through this path. The same pattern appears elsewhere. Model outputs are schema-validated and then checked against business rules that JSON Schema cannot conveniently express. A soundtrack must refer to a real entry from the indexed music library. The model cannot invent a filename and hope it exists. Human gates sit around the places where human judgment is actually worth the interruption, such as candidate approval and final publication. This looks different from the trading pipeline, but conceptually it is the same architecture. In the trading system, an LLM may propose code. Tests and deployment contracts determine whether that code can become an actuator. In the video system, an LLM may propose an edit. The renderer determines whether that edit can become pixels. **The model proposes. The actuator permits.** That is a much stronger safety mechanism than trying to construct a perfect prompt. ## My decision rule for agentic systems After building both systems, I use four questions when deciding how much autonomy to give a model: 1. **What does a wrong action cost?** Measure money, safety impact, reputation, data loss, and irreversibility, not how impressive the task sounds. 2. **Where the cost is high, who owns the route?** I prefer explicit software control flow, bounded model tasks, typed artifacts, tests, and gates before side effects. 3. **Where the blast radius is small, what judgment can I profitably delegate?** That is where agentic behavior becomes interesting. Let the model explore, select, retry, and compose instead of converting every fuzzy decision into hand-written branching logic. 4. **What is the real actuator?** Put the hard boundary there. Filesystem permissions, deployment APIs, renderers, database operations, payment calls, trading endpoints, and publish buttons deserve stronger constraints than the prompts upstream of them. This is why I am skeptical of architecture discussions that start with model capability. “Can the model do this?” is useful during a demo. It is not sufficient for a system design. A very capable model with the wrong control boundary is a liability. A less capable model inside a well-designed system can be surprisingly useful because failure is contained and observable. The best agent architecture is therefore not necessarily the one that maximizes autonomy. Sometimes good engineering means building a cage. Sometimes it means opening the cage and making the walls around the playground stronger instead. The stakes decide which one you need. The real question is not **“Can the agent do the task?”** It is **“Where do I put the control boundary?”** ## About the Author Igor Khokhriakov is a Principal Software Engineer with 17+ years of experience building complex software systems, including work at DESY in Hamburg, the Tango Controls consortium, the San Diego Supercomputer Center, and currently the HDF Group. He specializes in high-performance distributed systems, with deep expertise in Java, Kubernetes, cloud infrastructure, and observability. His systems have supported hundreds of experiments at large-scale research infrastructures including PETRA III and the ESS. His work has been published in peer-reviewed journals including SPIE Proceedings and the Journal of Synchrotron Radiation, and his systems continue to influence scientific software architecture today. More recently, Igor has been exploring the architecture of AI-native and agentic systems. His projects include a closed-loop algorithmic trading pipeline that turns natural-language ideas into tested and deployed AWS Lambda strategies, and an AI-assisted content production system in which models analyze footage, plan edits, select music, and construct short-form video while deterministic actuators enforce hard safety constraints. His current focus is on the engineering boundary between autonomous agents and deterministic workflows: how much control to give models, where to place guardrails, and how to design systems in which AI can make useful decisions without owning more of the runtime than the risk allows. ## Related Articles - [Never delegate the understanding](https://www.wearedevelopers.com/magazine/749-never-delegate-the-understanding) - [AI Eats the Verifiable First](https://www.wearedevelopers.com/magazine/765-ai-eats-the-verifiable-first) - [Vibe coding, creativity, craft and professionalism… are we making ourselves redundant?](https://www.wearedevelopers.com/magazine/579-vibe-coding-creativity-craft-and-professionalism-are-we-making-ourselves-redundant) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) ## Related Videos - [The day the chatbot asked for sudo](https://www.wearedevelopers.com/videos/100038-the-day-the-chatbot-asked-for-sudo) - [What Production Knows: Closing the Loop Between AI Agents and the Systems They Build](https://www.wearedevelopers.com/videos/100277-what-production-knows-closing-the-loop-between-ai-agents-and-the-systems-they-build) - [A True Story About Speeding Up the Wrong Things](https://www.wearedevelopers.com/videos/1930-a-true-story-about-speeding-up-the-wrong-things) - [5 things I wish I hadn’t done building my AI agent](https://www.wearedevelopers.com/videos/100145-5-things-i-wish-i-hadn-t-done-building-my-ai-agent) ## Related Jobs - [Senior AI Software Engineer iv)](https://www.wearedevelopers.com/jobs/ext/3175712-senior-ai-software-engineer-iv) at **Bosch-Gruppe Österreich** - [Staff Software Engineer, GitHub Intelligence (Copilot Agents)](https://www.wearedevelopers.com/jobs/ext/2650582-staff-software-engineer-github-intelligence-copilot-agents) at **GitHub** - [Senior AI Developer](https://www.wearedevelopers.com/jobs/ext/2836034-senior-ai-developer) at **PwC** - [Head of Agentic AI / Lead AI Engineer](https://www.wearedevelopers.com/jobs/48488-head-of-agentic-ai-lead-ai-engineer) at **1st solution consulting gmbh** - [Staff Software Engineer, Agentic Platform](https://www.wearedevelopers.com/jobs/48464-staff-software-engineer-agentic-platform) at **Docker, Inc.** - [Software Engineer III, Copilot Models Inference](https://www.wearedevelopers.com/jobs/ext/3332076-software-engineer-iii-copilot-models-inference) at **GitHub**