> Markdown version of [/jobs/ext/2645724-machine-learning-ops-engineer](https://www.wearedevelopers.com/jobs/ext/2645724-machine-learning-ops-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Machine Learning Ops Engineer - **Company:** Zone 5 Technologies - **Location:** United States (Remote available) - **Experience:** Experienced - **Salary:** $140,000.0 - $175,000.0 - **Contract:** Permanent contract - **Skills:** Microsoft Access, Application Programming Interfaces (APIs), Artificial Intelligence, User Authentication, Business Systems, Configuration Management, Continuous Integration, Information Engineering, Information Leak Prevention, Python (Programming Language), OAuth, Role-Based Access Control, Ansible, DataOps, Single Sign-On, Software Engineering, Management of Software Versions, AI Infrastructure, Data Logging, Enterprise Software Applications, Large Language Models, Prompt Engineering, Kubernetes, Infrastructure Automation Frameworks, Information Technology, Deployment Automation, Integration Frameworks, Data Management, Machine Learning Operations, Hardware Infrastructure, Data Pipelines - **Published:** August 4, 2026 - **Apply:** https://job-boards.greenhouse.io/zone5technologies/jobs/5381203008 ## About the Role * Bachelor's in Computer Science, Software Engineering, Data Engineering, or related field - equivalent industry experience also welcome * 3-6+ years of experience in MLOps, software, platform, or backend engineering (relevant depth matters more than exact years) * Strong proficiency in Python and comfort building, shipping, and operating services * Experience building LLM-powered applications-working with LLM APIs or self-hosted models, prompts, and tool/function calling * Hands-on experience with Kubernetes and containerized deployment * Solid understanding of CI/CD, infrastructure-as-code, and production service reliability * Awareness of access control and data-boundary concerns when connecting tools to sensitive internal systems * Demonstrated ability to learn quickly and work across unfamiliar parts of the stack * Depth in at least one core area-LLM application development, RAG/retrieval, agent design, or AI infrastructure-with genuine interest in growing into the others Preferred: * Hands-on experience with RAG systems, embeddings, and vector databases (pgvector, Qdrant, Weaviate, Milvus, or similar) * Experience designing and shipping agentic workflows, including tool use, orchestration, and guardrails * Familiarity with the Model Context Protocol (MCP) or similar tool-integration frameworks for LLMs * Experience integrating LLM tools with enterprise systems (productivity suites, business systems, or developer platforms) via their APIs * Knowledge of LLM evaluation, prompt engineering, and quality/regression measurement * Experience serving models and optimizing inference (vLLM, TGI, Triton, or similar) * Familiarity with agent/orchestration libraries (LangChain, LlamaIndex, or equivalent) * Experience with Ansible for configuration management and automation * Experience implementing identity, authentication, and fine-grained authorization (OAuth, SSO, RBAC) * Observability experience for AI/ML workloads, including usage and quality metrics * GPU infrastructure and scheduling experience for training or inference * Understanding of security and data-handling requirements in regulated or defense environments * Ability to obtain or maintain a security clearance ## Description LLM Applications, RAG & Agents * Design and build new LLM-powered tools and agentic workflows that automate real work and improve productivity across the company * Extend and improve our RAG systems-ingestion, chunking, embedding, retrieval, ranking, and evaluation-to raise answer quality * Structure retrieval around the organization's information hierarchy so that relevance and access boundaries improve together * Build tool integrations that connect LLMs to internal systems and data sources * Design agents that act safely against real systems, with appropriate guardrails, human-in-the-loop where warranted, and clear failure behavior * Establish evaluation and testing frameworks to measure quality, catch regressions, and guide iteration * Partner with teams across the company to identify high-value use cases and turn them into deployed tools Service Deployment & AI Infrastructure * Deploy AI tools and services for teams across the company, taking them from prototype to reliable production * Build and operate the infrastructure that hosts models, tools, and supporting services on Kubernetes * Manage model serving, inference endpoints, and the APIs and gateways around them * Implement monitoring, logging, and usage observability so we understand how tools perform and get used Access, Security & Data Boundaries * Ensure retrieval and agent tools respect the same access boundaries as the underlying systems-no cross-team or cross-project data leakage * Integrate with existing identity and permission systems so tools honor who is allowed to see what * Apply data-handling practices appropriate to a defense environment * Treat access control as a first-class design concern in every tool, not an afterthought Automation & Data Operations * Build CI/CD pipelines for AI tools, services, and agents * Automate provisioning and configuration with Ansible and infrastructure-as-code practices * Build data pipelines to ingest, transform, and index content for RAG and AI applications * Manage vector databases and other stores backing retrieval and AI workloads, including versioning and quality checks * Maintain reproducible environments across development, staging, and production ## Related Videos - [Keeping applications secure by evolving OAuth 2.0 and OpenID Connect](https://www.wearedevelopers.com/videos/100152-keeping-applications-secure-by-evolving-oauth-2-0-and-openid-connect) - [Dev & Test in the Cloud? Deploy your cloud environments with Ansible & Terraform](https://www.wearedevelopers.com/videos/1607-dev-test-in-the-cloud-deploy-your-cloud-environments-with-ansible-terraform) - [DevOps for AI: running LLMs in production with Kubernetes and KubeFlow](https://www.wearedevelopers.com/videos/1222-devops-for-ai-running-llms-in-production-with-kubernetes-and-kubeflow) - [Delay the AI Overlords: How OAuth and OpenFGA Can Keep Your AI Agents from Going Rogue](https://www.wearedevelopers.com/videos/1637-delay-the-ai-overlords-how-oauth-and-openfga-can-keep-your-ai-agents-from-going-rogue) - [Effective Machine Learning - Managing Complexity with MLOps](https://www.wearedevelopers.com/videos/185-effective-machine-learning-managing-complexity-with-mlops) - [Eclipse Che for Infrastructure Automation](https://www.wearedevelopers.com/videos/1611-eclipse-che-for-infrastructure-automation) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer)