> Markdown version of [/videos/2056-command-and-conquer-how-we-let-an-llm-control-our-software](https://www.wearedevelopers.com/videos/2056-command-and-conquer-how-we-let-an-llm-control-our-software). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Command and Conquer: How we let an LLM control our Software Simon Jimenez proves treating LLMs as unreliable APIs ensures stable software. Massive prompts fail. Discover how a plan-and-fulfill architecture safely gives AI the keys to your application. - **Speakers:** [Simon A.T. Jiménez](https://www.wearedevelopers.com/@simon-a-t-jimenez) - **Event:** World Congress 2026 Europe - Virtual Stage - **Published:** July 2, 2026 - **Duration:** 37:51 - **URL:** https://www.wearedevelopers.com/videos/2056-command-and-conquer-how-we-let-an-llm-control-our-software ## Summary Simon Jimenez, technical lead for the requirements management tool Storywise, explores the architectural journey of letting Large Language Models (LLMs) control software via a command pattern. Starting with a proprietary model in 2022 before migrating to an agnostic, multi-provider routing system (including Azure OpenAI, Anthropic, and European-hosted alternatives like Scaleway), the team tackled the inherent unpredictability of AI. By designing an LLM assistant that translates user inputs into structured commands, they created a powerful yet controlled interface for generating user stories and specifications. A major breakthrough came from transitioning to a "plan and fulfill" architecture. Rather than relying on a single, massive prompt burdened with complex JSON schemas, the system uses a lightweight planner to generate an operational to-do list, passing isolated, fresh contexts to a fulfiller step. Jimenez emphasizes treating LLMs as "extremely unreliable APIs," requiring developers to aggressively parse responses (such as stripping rogue markdown from JSON outputs) and gracefully degrade functionality so the core application remains fully operable offline. Furthermore, integrating robust token cost-tracking from day one is vital to prevent runaway API bills when users ingest massive datasets like full Confluence spaces. Crucially, the system strictly enforces a human-in-the-loop harness; the AI only proposes operations, preventing destructive actions without explicit user approval. When testing these inherently non-deterministic models, the team learned to validate the intent of the output rather than exact phrasing, accepting a "works most of the time" baseline. Coupled with per-tenant endpoint isolation and automated GDPR compliance reviews, this approach ensures enterprise-grade security and reliability, demonstrating that integrating conversational AI can "absolutely rock your software" if securely and thoughtfully implemented. **Keywords:** llm command pattern, multi-provider llm routing, ai structured output extraction, plan and fulfill architecture, human-in-the-loop ai, software requirements generation, llm token cost management, graceful ai degradation, gdpr compliance automation, enterprise llm hosting, json schema validation, non-deterministic ai testing, ai prompt versioning, vendor-agnostic ai integration, llm strict mode limitations ## Chapters 1. **Letting large language models control software applications** (00:00) — Implementing a command pattern allows users to naturalistically manage software requirements without needing manual inputs. 1. **Transitioning from proprietary models to multi-provider routing** (02:26) — Adopting an agnostic model infrastructure prevents vendor lock-in and accommodates regional hosting constraints. 1. **Handling provider inconsistencies and fallback offline modes** (06:45) — Building robust endpoint strategies manages varying API schema compliance and offline scenarios effectively. 1. **Mitigating language model unpredictability and parsing issues** (10:43) — Extracting structural output with a custom JSON parser handles the chattiness and language drifts of large models. 1. **Managing token consumption and knowledge base context** (14:27) — Implementing scoped context limits prevents unexpected invoice spikes from massive integrations like confluence document syncs. 1. **Securing data with tenant endpoints and handling timeouts** (19:43) — Allowing custom tenant endpoints prevents information leakage while asynchronous requests require careful timeout management. 1. **Designing safe user experiences with human approval** (22:34) — Proposing operations instead of auto-executing them ensures critical data modifications are explicitly reviewed by human operators. 1. **Improving reliability through plan and fulfill prompting** (26:39) — Splitting complex tasks into a lightweight planning phase and dedicated fulfillment steps drastically improves execution reliability. 1. **Automating compliance reviews with specialized context prompts** (31:21) — Running targeted system prompts against user stories enables automated GDPR checks and generates resolution suggestions. 1. **Key takeaways for treating language models as APIs** (33:32) — Treating large language models as unreliable external APIs demands strict validation and continuous prompt versioning. ## Related Moments - [Summarizing developer experience and artificial intelligence companions](https://www.wearedevelopers.com/videos/884-forget-developer-platforms-think-developer-productivity) (from "Forget Developer Platforms, Think Developer Productivity!") - [Enhancing conversational intent through modern large language models](https://www.wearedevelopers.com/videos/1641-hello-jarvis-building-voice-interfaces-for-your-llms) (from "Hello JARVIS - Building Voice Interfaces for Your LLMS") - [Transitioning from AI co-pilots to AI-native products](https://www.wearedevelopers.com/videos/100091-3-ways-to-rebuild-the-data-stack-for-agents) (from "3 Ways to Rebuild the Data Stack for Agents") - [Building components of a real-world LLM lifecycle](https://www.wearedevelopers.com/videos/1250-from-traction-to-production-maturing-your-llmops-step-by-step) (from "From Traction to Production: Maturing your LLMOps step by step") - [Introducing LLMs as judges for automated testing](https://www.wearedevelopers.com/videos/100300-testing-ai-agents-automated-evaluation-for-chatbots-rag-systems) (from "Testing AI Agents: Automated Evaluation for Chatbots & RAG Systems") - [Applying large language models to infrastructure tasks](https://www.wearedevelopers.com/videos/2084-your-infrastructure-is-not-a-playground-ai-agents-for-infra-done-right) (from "Your Infrastructure Is Not a Playground: AI Agents for Infra Done Right") ## Related Articles - [WWC24 Talk - Scott Hanselman - AI: Superhero or Supervillain?](https://www.wearedevelopers.com/magazine/469-wwc24-talk-scott-hanselman-ai-superhero-or-supervillain) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) ## Related Jobs - [Machine Learning Engineer](https://www.wearedevelopers.com/jobs/ext/588393-machine-learning-engineer) at **Twilio** - [Machine Learning Engineer](https://www.wearedevelopers.com/jobs/ext/1355348-machine-learning-engineer) at **TWILIO** - [Staff Software Engineer, Copilot Experiences](https://www.wearedevelopers.com/jobs/ext/164361-staff-software-engineer-copilot-experiences) at **GitHub** - [AI Software Engineer (Germany)](https://www.wearedevelopers.com/jobs/48317-ai-software-engineer-germany) at **Sunhat** - [Principal Software Engineer, Enterprise AI Platform](https://www.wearedevelopers.com/jobs/ext/1467292-principal-software-engineer-enterprise-ai-platform) at **GitHub** - [AI Operations Manager (all genders)](https://www.wearedevelopers.com/jobs/48263-ai-operations-manager-all-genders) at **envelio**