> Markdown version of [/jobs/ext/355326-ai-engineer-llm-infrastructure](https://www.wearedevelopers.com/jobs/ext/355326-ai-engineer-llm-infrastructure). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Engineer (LLM Infrastructure) - **Company:** Yoursafe - **Location:** Amsterdam, Netherlands - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Training Data, Application Programming Interfaces (APIs), Artificial Intelligence, Nvidia CUDA, Databases, Linux, Memory Management, Information Retrieval, Python (Programming Language), Open Source Technology, Search Technologies, AI Infrastructure, Data Ingestion, Pytorch, Large Language Models, Bare Metal, Data Analytics, Hardware Acceleration, Data Management, Machine Learning Operations, ZFS File System, TensorRT - **Published:** June 5, 2026 - **Apply:** https://nl.indeed.com/viewjob?jk=cb4fff9934ca4900 ## About the Role Do you have experience in Python?, Do you have a Master's degree?, · A senior AI/MLOps engineer with proven experience scaling self-hosted LLM infrastructure in a production environment. · Experienced with semantic search, vector databases, and information retrieval techniques (RAG) at scale. · Deeply experienced with Python, PyTorch, CUDA, and modern inference serving frameworks. · Experienced in using AI as a production and verification tool, not a gimmick. · Comfortable working closely with networking and storage architects to eliminate I/O bottlenecks in a ZFS/Linux ecosystem. · Highly structured, data-driven, and execution-focused. · Motivated by building systems that scale, not campaigns that win awards. ## Description Behind our scalable Open Issuing platform and our push into 100 target countries, sits a highly advanced, bare-metal AI infrastructure. We are building a high-density, on-premise AI and automation environment consisting of cutting-edge heavy compute (Nvidia H200), high-performance enterprise storage, and a fleet of automated agents., We are seeking a senior AI Engineer to take ownership of our LLM inference infrastructure. You will own the hardware and software running our AI models, ensuring that our internal automation agents and our growth teams have zero-latency, highly optimized access to state-of-the-art open-source LLMs. You will collaborate closely with our CEO, COO, Product Manager and Head of Growth to engineer the AI-driven pipelines required for large-scale content creation, localization and personalised agentic support. This is a senior, hands-on role. You are expected to manage hardware, deploy models, optimize inference speeds, and scale what works. Your responsibilities: LLM Deployment & Optimization · Own and manage our heavy-compute AI hardware, specifically optimizing workloads for our Nvidia HGX H200 infrastructure. · Deploy, fine-tune, and maintain open-source LLMs, ensuring maximum throughput and minimal latency. · Manage inference engines (e.g., vLLM, TensorRT-LLM) and handle dynamic GPU memory allocation for hundreds of concurrent agent requests. AI Automation & Agent Infrastructure · Provide a flawless, millisecond-response API layer for our "OpenClaw" agent farm (a fleet of 200+ bare-metal Apple Silicon nodes). · Monitor model performance, detect hallucinations, and build synthetic training data pipelines to continuously improve agent accuracy. · Design, scale, and maintain a high-performance, on-premise RAG service and vector database (e.g., Qdrant, Milvus, Milvus/Chroma) to seamlessly serve internal documentation to our LLMs. · Build robust data ingestion and embedding pipelines to ensure internal knowledge bases and documents are updated in real-time for the RAG service. Cross-Functional AI Tooling · Work tightly with Product, Operations, and the Growth team to ensure alignment. · Provide the technical foundation for the Growth team's AI-assisted content creation, verification, and contextual validation pipelines. · Build repeatable AI engines that can guarantee linguistic, cultural, and regulatory correctness across dozens of markets simultaneously. ## Related Videos - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [DevOps for AI: running LLMs in production with Kubernetes and KubeFlow](https://www.wearedevelopers.com/videos/1222-devops-for-ai-running-llms-in-production-with-kubernetes-and-kubeflow) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Dev Digest 132 - Binging WADFlix?](https://www.wearedevelopers.com/magazine/473-dev-digest-132-binging-wadflix)