> Markdown version of [/jobs/ext/2719933-software-engineer-agentic-systems](https://www.wearedevelopers.com/jobs/ext/2719933-software-engineer-agentic-systems). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer, Agentic Systems - **Company:** HORIZON 3, LLC - **Location:** Chicago, IL, United States (Remote available) - **Experience:** Expert - **Salary:** $169,000.0 - $208,000.0 - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Software System Penetration Testing, Python (Programming Language), Neo4j, Web Application Security, SQL Injection, Web Applications, Large Language Models, Multi-Agent Systems, Cross-Site Scripting (XSS) - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/senior-software-engineer-agentic-systems-horizon3-ai-8400257 ## About the Role * 5+ years building production software, with strong Python. * Hands-on experience building LLM-powered applications or agents, tool use / function calling, structured outputs, multi-step orchestration, and the glue that makes it all hold together. * A track record of making LLMs reliable in production, you've wrestled nondeterminism, designed around model limitations, and shipped something that worked when it mattered. * Real experience with evaluation: you've built or owned the harness that tells you whether a model or agent change is an improvement, not just a vibe. * Strong instincts for prompt and context engineering, and the judgment to keep the model's job small and well-scoped. * Solid software fundamentals - testing, observability, and the discipline to keep a complex agent debuggable. * Ownership mentality, comfortable owning a critical, fast-moving subsystem end to end. Desired/Nice to Have * Working knowledge of web application security, broken access control, IDOR/BOLA, SQLi, XSS, SSRF, SSTI, enough to collaborate fluently with offensive engineers. * Experience building eval harnesses or benchmarks specifically for agents (synthetic environments, CVE-based test targets, capture-the-flag-style scoring). * Experience with agent frameworks, and strong opinions about when not to reach for one. * Familiarity with graph data models (e.g., Neo4j) for representing application state and attack context. What makes you stand out: * You've shipped an autonomous agent that did real, valuable work unattended in production, and you have scar tissue from making it trustworthy. * You've designed evaluation systems that actually drove improvement, closed the loop between "we changed something" and "it measurably got better." * You pair an offensive-security mindset (CTF, bug bounty, pentesting, or research background) with the engineering chops to turn that intuition into a reliable system. * You have hands-on experience with agent fine-tuning or RL (SFT, GRPO, reward design for tool-using agents) and a grounded view of when it's worth it versus improving the harness. * You've published or spoken on agent reliability, evaluation, or autonomous security tooling. ## Description We're building an autonomous, black-box web application penetration tester. It crawls and attacks real production websites the way a skilled human pentester would, finding broken access control, injection, XSS, SSRF, SSTI, and more, under a strict production-safe, no-false-positives mandate., * Build and evolve the agent harness and orchestration that turns an LLM into a reliable autonomous pentester, the loop that reasons over an application, forms attack hypotheses, acts, and verifies results. * Design the tools and tool-shaped feedback the agent uses to probe and exploit, and the structured-output and validation layers that keep it reliable (e.g., hook-enforced mandatory validation, schema-constrained outputs). * Translate the team's offensive expertise into repeatable agent capabilities - partnering directly with our attackers to encode how they think into something the agent can do consistently. * Own and grow our evaluation infrastructure: benchmark suites, a failure-mode taxonomy across the pipeline (discovery * hypothesis * exploitation * verification), and regression detection, so we actually know whether the agent is getting better. * Manage LLM inference in production: model selection, prompt and context engineering, and keeping cost and latency under control (we run on AWS Bedrock with centralized cost tracking). * Hold the line on production-safety and no-false-positives, every finding the agent reports has to be real and reproducible. ## Related Videos - [Putting the Graph In GraphQL With The Neo4j GraphQL Library](https://www.wearedevelopers.com/videos/257-putting-the-graph-in-graphql-with-the-neo4j-graphql-library) - [Hacking MSSQL on Cloud. All of them. How I became sysadmin on Azure, AWS, GCP and Alibaba.](https://www.wearedevelopers.com/videos/100339-hacking-mssql-on-cloud-all-of-them-how-i-became-sysadmin-on-azure-aws-gcp-and-alibaba) - [Generate AI in the Browser with Chrome AI - Raymond Camden](https://www.wearedevelopers.com/videos/1770-generate-ai-in-the-browser-with-chrome-ai-raymond-camden) - [Agentic AI: building autonomous systems for developers](https://www.wearedevelopers.com/videos/100034-agentic-ai-building-autonomous-systems-for-developers) - [Cyber Sleuth: Finding Hidden Connections in Cyber Data](https://www.wearedevelopers.com/videos/893-cyber-sleuth-finding-hidden-connections-in-cyber-data) - [WebXR: Enabling Virtual and Augmented Reality on the Web](https://www.wearedevelopers.com/videos/965-webxr-enabling-virtual-and-augmented-reality-on-the-web) ## Related Articles - [What is Agentic Programming and Why Should Developers Care?](https://www.wearedevelopers.com/magazine/625-what-is-agentic-programming-and-why-should-developers-care) - [A 5-Step Open-Source Setup for Agentic Engineering](https://www.wearedevelopers.com/magazine/738-a-5-step-open-source-setup-for-agentic-engineering) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path) - [Never delegate the understanding](https://www.wearedevelopers.com/magazine/749-never-delegate-the-understanding) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again)