> Markdown version of [/jobs/ext/3029721-ai-os-engineer](https://www.wearedevelopers.com/jobs/ext/3029721-ai-os-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI OS Engineer - **Company:** Harvey Nash - **Location:** Dallas, TX, United States (Remote available) - **Experience:** Experienced - **Salary:** $166,400.0 - $208,000.0 - **Contract:** Temporary to permanent - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Amazon Web Services, Microsoft Azure, Python (Programming Language), Queueing Systems, RabbitMQ, Redis, Azure Machine Learning, Systems Architecture, Cloud Platform System, Large Language Models, Multi-Agent Systems, Containerization, Kubernetes, Apache Kafka, Virtual Agents, Restful APIs, Docker, Microservices - **Published:** September 22, 2026 - **Apply:** https://careers.harveynashusa.com/jobsearch/job-details/ai-os-engineer-8426-1/ ## About the Role * 5+ years of production software engineering experience, with 2+ years focused on building agentic frameworks, multi-agent orchestrations, or LLM infrastructure. * Advanced proficiency in Python, TypeScript/Node.js, and modern async execution models. * Hands-on expertise with agent architectures and orchestration frameworks (e.g., LangGraph, AutoGen, CrewAI, LlamaIndex, or custom in-house runtimes). * Proven track record working with vector databases, embedding systems, and hybrid RAG implementations. * Direct experience with Docker, Kubernetes, vLLM / Triton inference engines, and cloud platforms (AWS Sagemaker, GCP Vertex AI, or Azure ML). * Mastery of RESTful/gRPC APIs, message queues (Kafka, RabbitMQ, Redis), and microservice architectures., * Experience with local LLM serving, quantization methods (AWQ, GGUF), and self-hosted foundation models (Llama, Mistral). * Deep understanding of sandboxed execution environments (e.g., WebAssembly, Docker-in-Docker, E2B) for safe AI agent tool execution. * Prior contract experience operating in fast-paced, 6-month delivery cycles with clear milestone check-ins. ## Description We are seeking an experienced AI OS Engineer for a high-impact, 6-month contract initiative. In this role, you will lead the architecture and integration of our next-generation AI Operating System (AI OS)-a core orchestration framework designed to seamlessly manage autonomous agents, multi-LLM routing, context memory systems, tool execution, and local-to-cloud compute pipelines., * Design, build, and deploy agentic workflows, dynamic task schedulers, and execution runtime environments powering internal AI applications. * Implement robust retrieval systems, long-term state persistence, vector databases (e.g., pgvector, Qdrant, Pinecone), and hybrid-search mechanisms to optimize agent context windows. * Architect multi-model routing layers (e.g., Anthropic, OpenAI, open-source foundation models) for cost-efficiency, fallback management, and low-latency inference. * Develop secure sandbox environments for tool execution, code generation, API calls, and agent safety protocols. * Build evaluation harnesses to track model drift, execution accuracy, hallucination rates, and latency bottlenecks. * Containerize and deploy AI OS infrastructure on cloud environments (AWS / GCP / Azure) using CI/CD pipelines., * Finalize system architecture, set up local/cloud runtime execution environments, and deploy the core orchestration layer. * Integrate multi-agent tool execution, long-term memory state persistence, and guardrail protocols. * Conduct system-wide evaluation harness benchmarking, latency/cost optimization, and handoff documentation for internal engineering teams. ## Related Videos - [Agentic AI - From Theory to Practice: Developing Multi-Agent AI Systems on Azure](https://www.wearedevelopers.com/videos/1532-agentic-ai-from-theory-to-practice-developing-multi-agent-ai-systems-on-azure) - [Beyond Kafka & RabbitMQ: Why NATS is the Future of Microservices Messaging](https://www.wearedevelopers.com/videos/1646-beyond-kafka-rabbitmq-why-nats-is-the-future-of-microservices-messaging) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Reducing LLM Calls with Vector Search Patterns - Raphael De Lio (Redis)](https://www.wearedevelopers.com/videos/1714-reducing-llm-calls-with-vector-search-patterns-raphael-de-lio-redis) - [Developing ASP.NET Core Microservices with Dapr: A practical guide](https://www.wearedevelopers.com/videos/1528-developing-asp-net-core-microservices-with-dapr-a-practical-guide) - [Why Systems Break After Initial Success: The Architectural Failures That Take Months to Surface](https://www.wearedevelopers.com/videos/2048-why-systems-break-after-initial-success-the-architectural-failures-that-take-months-to-surface) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [What is Agentic Programming and Why Should Developers Care?](https://www.wearedevelopers.com/magazine/625-what-is-agentic-programming-and-why-should-developers-care) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift)