AI OS Engineer

Harvey Nash
Dallas, TX, United States
10 days ago
Apply on careers.harveynashusa.com
Prepare application

Role details

Contract type
Temporary to permanent
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
2 years minimum
Compensation
$166,400.0 - $208,000.0
Working hours
Regular working hours

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Amazon Web Services Microsoft Azure Python (Programming Language) Queueing Systems RabbitMQ Redis Azure Machine Learning Systems Architecture Cloud Platform System Large Language Models
+8 more
Multi-Agent Systems Containerization Kubernetes Apache Kafka Virtual Agents Restful APIs Docker Microservices

Job description

We are seeking an experienced AI OS Engineer for a high-impact, 6-month contract initiative. In this role, you will lead the architecture and integration of our next-generation AI Operating System (AI OS)-a core orchestration framework designed to seamlessly manage autonomous agents, multi-LLM routing, context memory systems, tool execution, and local-to-cloud compute pipelines., * Design, build, and deploy agentic workflows, dynamic task schedulers, and execution runtime environments powering internal AI applications.

  • Implement robust retrieval systems, long-term state persistence, vector databases (e.g., pgvector, Qdrant, Pinecone), and hybrid-search mechanisms to optimize agent context windows.
  • Architect multi-model routing layers (e.g., Anthropic, OpenAI, open-source foundation models) for cost-efficiency, fallback management, and low-latency inference.
  • Develop secure sandbox environments for tool execution, code generation, API calls, and agent safety protocols.
  • Build evaluation harnesses to track model drift, execution accuracy, hallucination rates, and latency bottlenecks.
  • Containerize and deploy AI OS infrastructure on cloud environments (AWS / GCP / Azure) using CI/CD pipelines., * Finalize system architecture, set up local/cloud runtime execution environments, and deploy the core orchestration layer.
  • Integrate multi-agent tool execution, long-term memory state persistence, and guardrail protocols.
  • Conduct system-wide evaluation harness benchmarking, latency/cost optimization, and handoff documentation for internal engineering teams.

Requirements

  • 5+ years of production software engineering experience, with 2+ years focused on building agentic frameworks, multi-agent orchestrations, or LLM infrastructure.
  • Advanced proficiency in Python, TypeScript/Node.js, and modern async execution models.
  • Hands-on expertise with agent architectures and orchestration frameworks (e.g., LangGraph, AutoGen, CrewAI, LlamaIndex, or custom in-house runtimes).
  • Proven track record working with vector databases, embedding systems, and hybrid RAG implementations.
  • Direct experience with Docker, Kubernetes, vLLM / Triton inference engines, and cloud platforms (AWS Sagemaker, GCP Vertex AI, or Azure ML).
  • Mastery of RESTful/gRPC APIs, message queues (Kafka, RabbitMQ, Redis), and microservice architectures., * Experience with local LLM serving, quantization methods (AWQ, GGUF), and self-hosted foundation models (Llama, Mistral).
  • Deep understanding of sandboxed execution environments (e.g., WebAssembly, Docker-in-Docker, E2B) for safe AI agent tool execution.
  • Prior contract experience operating in fast-paced, 6-month delivery cycles with clear milestone check-ins.

Benefits & conditions

Medical, dental, and vision coverage 401(k) retirement plan Voluntary benefits and insurance options Referral bonus opportunities Pre-tax commuter benefits

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on careers.harveynashusa.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

4:37 min

Transitioning to a hybrid human and agent workforce

Thomas Dohmke Thomas Dohmke · World Congress 2025

3:55 min

Demonstrating semantic routing thresholds with the Redis vector library

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

49 sec

Processing asynchronous event streams via standard messaging patterns

Patrick Koss Patrick Koss · World Congress 2024

3:45 min

Fusing developer experience and platform engineering for agentic SDLC

Julia Kordick Julia Kordick · World Congress 2026 Europe

3:42 min

Comparing in-memory and Redis storage for cache scalability

Simone Sanfratello · World Congress 2022

Videos

See all

Related articles

See all