> Markdown version of [/jobs/ext/348709-software-inference-deployment-engineer](https://www.wearedevelopers.com/jobs/ext/348709-software-inference-deployment-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Inference Deployment Engineer - **Company:** Lumai - **Location:** Oxford, UK - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Computing Platforms, Nvidia CUDA, Data Centers, Linux, Field-Programmable Gate Array (FPGA), Python (Programming Language), Tensorflow, Software Deployment, Software Engineering, Systems Integration, AI Infrastructure, Graphics Processing Unit (GPU), Pytorch, Large Language Models, Software Troubleshooting, Containerization, Machine Learning Operations, TensorRT, Docker - **Published:** June 19, 2026 - **Apply:** https://uk.indeed.com/viewjob?jk=f15999c70daa36fa ## About the Role Do you have experience in Software deployment?, Do you have a Master's degree?, Must-Have * Hands-on software engineering experience in AI infrastructure, inference serving, accelerator integration, or comparable deep-tech hardware-software environments * Strong Python skills and familiarity with major ML frameworks (PyTorch in particular) * Practical experience with model deployment workflows - loading, format conversion, quantisation, or framework integration * Comfortable working with inference serving stacks (for example vLLM, TensorRT-LLM, or similar) * Familiarity with Linux, containerisation (Docker), and cluster environments * Comfortable in a customer-facing role, able to communicate clearly with ML and infrastructure engineering teams * Comfortable working in a fast-moving, early-stage environment where the product and the deployment approach are both still being developed Strong Preference For * Experience integrating accelerator hardware (GPUs, FPGAs, ASICs, NPUs, or novel architectures) into customer inference workflows * Familiarity with the NVIDIA inference stack - CUDA, TensorRT, Triton * Exposure to disaggregated inference architectures, prefill/decode separation, or KV cache management ## Description We are bringing the world's first optical AI compute platform to market. As we move from development into field deployment, we are looking for a Software Inference Deployment Engineer to own the software-side integration and customer support of Lumai Iris servers in third-party data centre environments. You will begin by working alongside our software and engineering teams - helping integrate the Iris software stack, supporting model onboarding through the toolchain, and getting hands-on with the disaggregated prefill/decode runtime. This is intentional: the best way to develop deep expertise in a novel platform is to build with it. As deployments go live, you will take ownership in the field - supporting customer integration into their inference stacks, troubleshooting software issues, and acting as a primary technical contact for customer ML and infrastructure engineering teams. This is an opportunity to work at the cutting edge of efficient AI inference - deploying a genuinely novel compute platform into production for the first time, and playing a central role in how it reaches the world., * Work alongside Lumai's software and engineering teams to integrate, test, and harden the Iris software stack ahead of deployment * Support model onboarding through the Iris toolchain - loading, conversion, and framework integration * Develop hands-on familiarity with the disaggregated prefill/decode runtime, including how Iris servers operate alongside decode processors * Support customer integration of Lumai Iris into their own frameworks * Own software-side troubleshooting in the field, acting as the first line of response post-deployment * Train and enable customer ML and infrastructure engineering teams on the Iris software platform * Feed field findings, integration issues, and customer feedback back into product and engineering ## Related Videos - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [LLMOps-driven fine-tuning, evaluation, and inference with NVIDIA NIM & NeMo Microservices](https://www.wearedevelopers.com/videos/1582-llmops-driven-fine-tuning-evaluation-and-inference-with-nvidia-nim-nemo-microservices) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud)