> Markdown version of [/jobs/ext/2546661-ai-infrastructure-engineer-agents-ml-systems](https://www.wearedevelopers.com/jobs/ext/2546661-ai-infrastructure-engineer-agents-ml-systems). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Infrastructure Engineer - Agents & ML Systems - **Company:** HavocAI Inc - **Location:** United States (Remote available) - **Experience:** Experienced - **Salary:** $175,000.0 - $200,000.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Airflow, Amazon S3, Data Analysis, Systems Engineering, Automated Storage and Retrieval Systems, Audit Trail, C++ (Programming Language), Computer Programming, Databases, Continuous Integration, Data Cleansing, Information Engineering, Data Files, Data Systems, Software Debugging, Programming Tools, Document Management Systems, Distributed Systems, Github, Graph Database, Python (Programming Language), Key Management, PostgreSQL, Machine Learning, Regression Testing, Search Technologies, Software Engineering, Systems Integration, TypeScript, AI Infrastructure, Data Logging, Software Repository, Large Language Models, Model Validation, Backend, Data Lakes, Kubernetes, Information Technology, Apache Kafka, Build Tools, Machine Learning Operations, Data Pipelines, Docker, Golang, Data Generation - **Published:** August 4, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=e4f61deb16448fb4 ## About the Role * Bachelor's degree in Computer Science, Engineering, Machine Learning, Data Science, Applied Mathematics, or a related technical field. * 3+ years of experience in software engineering, infrastructure engineering, ML infrastructure, backend systems, data engineering, developer tools, or related technical roles. * Strong programming experience in Python, TypeScript, Go, C++, or similar languages. * Experience building production software systems, APIs, services, data pipelines, or internal platforms. * Experience with, or strong interest in, LLM applications, AI agents, tool-using systems, RAG pipelines, or AI developer tools. * Familiarity with modern AI infrastructure concepts such as embeddings, vector search, prompt management, evaluation, model serving, fine-tuning, or MLOps. * Ability to integrate systems across APIs, databases, object stores, documents, logs, internal tools, and structured or unstructured data sources. * Strong understanding of production engineering fundamentals, including reliability, observability, testing, and maintainability. * Strong grounding in securing AI and agentic systems, including least-privilege tool access, prompt-injection and misuse mitigation, secrets management, and safe handling of sensitive data. * Ability to work across Software, Data, ML, Infrastructure, and Product teams. * Strong debugging skills and comfort working with complex distributed systems. * U.S. citizenship and the ability to obtain and maintain a security clearance. Preferred Skills * Experience with MCP, including MCP servers, clients, tools, resources, prompts, or connector patterns. * Experience with LangGraph, LangChain, LlamaIndex, Semantic Kernel, OpenAI APIs, Anthropic APIs, local/open-weight models, or similar AI tooling. * Experience with vector databases, embeddings, retrieval systems, knowledge graphs, document processing, search, or RAG systems. * Experience with LLM fine-tuning, supervised fine-tuning, preference tuning, synthetic data generation, evaluation datasets, or model benchmarking. * Experience with Kubernetes, Docker, Ray, Airflow, Dagster, MLflow, Weights & Biases, Kafka, Postgres, S3-compatible storage, or similar infrastructure. * Experience building internal platforms, developer tools, workflow automation systems, or data/ML infrastructure. * Experience with human-in-the-loop workflows, approval systems, audit logs, policy enforcement, or safe tool invocation patterns. * Experience integrating AI tools with engineering workflows such as GitHub, CI/CD, issue trackers, documentation systems, simulation platforms, or data lakes. * Familiarity with autonomy, robotics, simulation, telemetry, perception, or defense technology workflows. * Experience building secure AI systems for sensitive, regulated, government, defense, or enterprise environments. * Active or prior security clearance. ## Description As an AI Infrastructure Engineer - Agents & ML Systems, you will help build the internal AI infrastructure that allows HavocAI teams to use modern AI systems safely, reliably, and effectively. You will develop tools, services, pipelines, and integrations that connect large language models, agentic workflows, internal data sources, engineering systems, and ML workflows. This role is ideal for a strong software or infrastructure engineer who is excited about the practical application of AI. You need not have worked on every part of the AI stack, but you should be curious, hands-on, and comfortable building production systems that connect models, tools, data, and users. You will work on systems that help internal teams search and reason over company data, automate engineering workflows, support simulation and autonomy development, curate data for future model training, and evaluate AI systems before they are trusted in critical workflows. This is a high-impact role at the intersection of software engineering, AI infrastructure, developer tooling, data systems, and applied ML., * Build internal AI infrastructure that connects LLMs and AI agents with internal tools, APIs, data sources, data lakes, telemetry stores, simulation tools, code repositories, documentation systems, logs, and engineering workflows. * Develop and maintain agentic AI systems for task automation, data analysis, engineering support, simulation workflows, and internal productivity. * Build tool integration and connector infrastructure for AI agents, including MCP and other emerging tool-use standards, spanning servers, tools, resources, prompts, connectors, and secure tool-use patterns. * Create pipelines for retrieval, RAG, context management, document processing, embeddings, and internal knowledge search. * Support ML infrastructure workflows such as data preparation, dataset curation, experiment tracking, model evaluation, fine-tuning support, and model deployment. * Build evaluation frameworks for agent performance, tool-use reliability, task success, model quality, regression testing, and failure analysis. * Develop observability, logging, tracing, auditability, monitoring, and debugging tools for AI agents, model calls, MCP tools, and ML pipelines. * Partner with Autonomy, Software, Data, Simulation, Product, and Operations teams to identify high-value AI use cases and turn them into reliable internal tools. * Secure agentic AI systems end-to-end with least-privilege tool access, sandboxed tool execution, prompt-injection and misuse mitigation, secrets management, human-in-the-loop approvals, and safe handling of sensitive and defense data. * Maintain documentation, reusable examples, templates, and best practices that help internal teams adopt AI tools safely and effectively. ## Related Videos - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [DevOps for AI: running LLMs in production with Kubernetes and KubeFlow](https://www.wearedevelopers.com/videos/1222-devops-for-ai-running-llms-in-production-with-kubernetes-and-kubeflow) - [Scoring 2000 Products per Request: Performance Pitfalls in Golang](https://www.wearedevelopers.com/videos/2073-scoring-2000-products-per-request-performance-pitfalls-in-golang) - [Bringing AI Model Testing and Prompt Management to Your Codebase with GitHub Models](https://www.wearedevelopers.com/videos/1536-bringing-ai-model-testing-and-prompt-management-to-your-codebase-with-github-models) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Dev Digest 210: AI Agents Are Go! Is MCP Dead? LLMs Crack Anonymity](https://www.wearedevelopers.com/magazine/709-dev-digest-210-ai-agents-are-go-is-mcp-dead-llms-crack-anonymity)