> Markdown version of [/jobs/ext/2864951-ai-agent-engineer-cloud-infrastructure](https://www.wearedevelopers.com/jobs/ext/2864951-ai-agent-engineer-cloud-infrastructure). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Agent Engineer, Cloud Infrastructure - **Company:** Orbis Operations - **Location:** McLean, VA, United States (Remote available) - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Amazon Web Services, Microsoft Azure, Catalyst (Software), Cloud Computing, Continuous Integration, Data Governance, Data Retrieval, Cursor (Graphical User Interface Elements), Software Debugging, Monitoring of Systems, Intelligence Analysis, JSON, Python (Programming Language), Key Management, Open Source Intelligence, Zero Trust Network Access, Runbook, Software Engineering, Website Wireframe, Datadog, Data Logging, Pulumi, Google Cloud, Cloud Platform System, Large Language Models, Multi-Agent Systems, Prompt Engineering, Multi-Cloud, Kubernetes, Cloudflare, Virtual Agents, GPT, Serverless Computing - **Published:** September 12, 2026 - **Apply:** https://diversityjobs.com/main/sendform/8/8/28176/1/18274583?backUrl=%2Fcareer%2F18274583%2FAi-Agent-Engineer-Cloud-Infrastructure-Virginia-Mclean ## About the Role Orbis weighs what a candidate has built over tenure or credentials. The following are true minimum requirements for this role: * Must have 5 years of proven experience translating user requirements into wireframes, workflows and such. * Must have 1-year proven experience utilizing Claude or similar commercial offering to build AI agents. * Demonstrated experience building a functioning multi-agent workflow, with the ability to walk through the design rationale, how the agents hand off work and share context, what broke, and how it was fixed * Practical command of prompt engineering, including retrieval-augmented generation (RAG) and tool or function calling * Proven ability to explain technical work to non-technical audiences - describing a system's behavior, limits, and failure modes to a process owner or executive without hedging or retreating into jargon * Demonstrated ability to elicit requirements from stakeholders who cannot hand over a specification, and to reach a working solution regardless * Ability to troubleshoot agent workflows built on platforms such as Claude, GPT-class models, LangChain, AutoGen, or CrewAI, and to isolate whether a failure originates in the prompt, the tool, the data, or the orchestration * Working proficiency with APIs and sufficient scripting ability to build and debug workflows independently; software engineering experience is not required, but a JSON payload or a Python script must not be a blocker * Sound judgment about model output, including the discipline to state plainly when an output should not be trusted * Strong written communication. This role produces documentation, runbooks, and reports that program leads and executives rely on * Hands-on production experience with at least one major cloud platform (GCP, Azure, or Cloudflare) and working familiarity with a second * Practical experience managing infrastructure as code using Terraform, Pulumi, or an equivalent * Working command of cloud networking, identity, and secrets management, and the ability to treat security posture and run cost as design constraints rather than afterthoughts Desired Qualifications Strong candidates will bring depth in one or more of the following. Do not screen yourself out for lacking items on this list. * 5-7 years experience in software development * Experience coding using Python * Cloud depth: Kubernetes or serverless runtimes at production scale; multi-cloud or cross-boundary deployment; cloud migration or consolidation work * Cost and governance: FinOps or formal cloud cost management practice; zero-trust architectures; federated access control models * Delivery and observability: owning CI/CD pipelines end to end; infrastructure monitoring, logging, and alerting for production workloads * Bridging technical and operational teams: prior work as a solutions consultant, technical program manager, forward-deployed engineer, or in a product-adjacent role * AI reliability practice: monitoring, evaluation, or observability tooling for AI systems in production; automated evaluations or regression checks for LLM-based workflows; human-in-the-loop review steps that maintain quality without creating bottlenecks * Process design: designing SOPs or escalation frameworks that hold up across multiple stakeholders * AI-assisted development tooling: working fluency with Claude, Copilot, or Cursor * Mission environment: operating within or supporting a government or defense program; OSINT, DIGINT, or intelligence analysis workflows and tradecraft; data governance concepts, federated access control models, or zero-trust architectures * Operating range: a track record of coming up to speed quickly on unfamiliar systems, and effectiveness in a distributed, multi-program environment with competing priorities Physical Requirements * Prolonged periods of sitting at a desk and working on a computer * Participation in virtual and in-person meetings * Ability to attend planned meetings and/or work in classified spaces for extended periods within the specified work regions * Travel: Occasional (up to 10%) to McLean, VA or team offsites as needed ## Description Orbis is building the agentic AI capability its mission support operations run on, and the AI Agent Engineer builds and operates it. Sixty percent of this role is common to every AI Agent Engineer at Orbis: designing, building, and deploying the multi-agent workflows behind triage, data retrieval, summarization, reporting, and escalation routing. The remaining forty percent is this role's focus - the cloud platform those workflows run on. This engineer owns the footprint across GCP, Cloudflare, and Azure, manages it as code, and holds its networking, security posture, and run cost to a standard that survives scrutiny as volume grows., Every AI Agent Engineer at Orbis, regardless of focus area, owns the following: * Design, build, and deploy agentic AI workflows. Architect and ship the multi-agent and automated workflows that run in production: triage, data retrieval, summarization, reporting, and escalation routing. * Translate in both directions. Explain to non-technical stakeholders what the systems do, where they are reliable, and where they are not, in terms those stakeholders can act on. Equally, take a process owner's plain-language description of how their work happens and turn it into a workflow specification that holds up. * Partner with process owners to build what they actually need. Work directly and iteratively with the colleagues who own a process so the resulting capability fits their operation rather than forcing them to fit it. Push back constructively when a stated requirement will not survive contact with production. * Keep the workflows reliable. Monitor what is running in production, diagnose failures across prompts, tools, data, and orchestration, and resolve them. Know which problems are yours to solve and when to bring in engineering. * Own the cost of execution. Track what your AI workloads cost to run. Identify where a tighter prompt, a caching layer, or a simple deterministic step delivers the same outcome, and keep those costs defensible as volume grows. * Set the quality bar for AI-assisted output. Define accuracy standards, human review checkpoints, and escalation criteria for the automated workflows, and build the checks that prove those standards are being met. * Document what you build. Author and maintain the runbooks and SOPs that allow others to operate and troubleshoot these workflows without you. Nothing lives only in one person's head. * Integrate across Orbis products. Work across Catalyst, Pulse, and Discovery, with engineering support where needed, so workflows can access and act on the data they require. * Deliver the reporting leadership and the team depend on. Build and maintain the outputs behind resolution rate, escalation rate, output accuracy, response time, and employee satisfaction, and turn the Senior Operations Manager's requirements into structured, recurring reports and data exports. * Close capability gaps and keep the practice current. Recognize when a need is not served by existing tooling, scope the gap clearly, and either build it or bring it to engineering and product with enough detail to act on. Track developments in agentic AI tooling and enterprise automation platforms, and form a defensible view on what is worth adopting, at what pace, and for which use cases. Focus Area - Cloud Infrastructure (40% of the Role) In addition to the common core above, this role owns the cloud infrastructure Orbis depends on - both the enterprise platform behind internal operations and the environments that support client delivery: * Design, deploy, and maintain the Orbis cloud footprint across GCP, AWS, Cloudflare, and Azure, spanning both enterprise workloads and client delivery environments. * Define and manage that footprint as infrastructure as code so environments are reproducible, reviewable, and recoverable. * Own networking, identity, and secrets handling for agent workloads, including the boundaries between environments and the audit of access. * Hold the security posture of the platform and keep it evidenced rather than asserted. * Track and control platform run cost across compute, inference, storage, and egress, and report it in terms leadership can act on. * Build and maintain the deployment and release path - CI/CD, environment promotion, rollback - so shipping an agent workflow is routine rather than an event. ## Related Videos - [On a Secret Mission: Developing AI Agents](https://www.wearedevelopers.com/videos/1510-on-a-secret-mission-developing-ai-agents) - [Why segmenting your infrastructure into tiers makes your infrastructure design better](https://www.wearedevelopers.com/videos/1960-why-segmenting-your-infrastructure-into-tiers-makes-your-infrastructure-design-better) - [Tips and Tricks for Working with JSON](https://www.wearedevelopers.com/videos/1229-tips-and-tricks-for-working-with-json) - [ Evaluating AI models for code comprehension](https://www.wearedevelopers.com/videos/1462-evaluating-ai-models-for-code-comprehension) - [Unleashing Potential Across Teams: The Power of Infrastructure as Code](https://www.wearedevelopers.com/videos/930-unleashing-potential-across-teams-the-power-of-infrastructure-as-code) - [Introducing JSON Structure](https://www.wearedevelopers.com/videos/100219-introducing-json-structure) ## Related Articles - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path) - [A 5-Step Open-Source Setup for Agentic Engineering](https://www.wearedevelopers.com/magazine/738-a-5-step-open-source-setup-for-agentic-engineering) - [What is Agentic Programming and Why Should Developers Care?](https://www.wearedevelopers.com/magazine/625-what-is-agentic-programming-and-why-should-developers-care) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [Prompt Engineering is a Job of the Past](https://www.wearedevelopers.com/magazine/342-prompt-engineering-is-a-job-of-the-past) - [The Prompt Engineer ✍️](https://www.wearedevelopers.com/magazine/216-the-prompt-engineer)