> Markdown version of [/jobs/ext/1368508-ai-inference-engineer](https://www.wearedevelopers.com/jobs/ext/1368508-ai-inference-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Inference Engineer - **Company:** Tether Operations Limited - **Location:** Barcelona, Spain (Remote available) - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Computer Vision, C++ (Programming Language), Nvidia CUDA, Computer Programming, Python (Programming Language), Machine Learning, Performance Tuning, Tensorflow, Pytorch, Large Language Models, Information Technology, ONNX (Open Neural Network Exchange) Format - **Published:** July 21, 2026 - **Apply:** https://www.adzuna.es/contact-us.html ## About the Role + Excellent programming skills in Python, and a solid understanding of C/C++. + Experience with platforms such as Llama.cpp, ONNX, TVM, MLC LLM, and IREE (MLIR), which facilitate the deployment of models to specific GPU architectures. + Experience in NLP, transformers, fine-tuning, computer vision, TensorFlow, PyTorch, JAX and CUDA toolkit. + Experience working with LLMs, fine tuning, RAG, transformers is a plus. + Demonstrated ability to rapidly assimilate new technologies and techniques. + A degree in Computer Science, AI, Machine Learning, or a related field, complemented by a solid track record in AI R&D. ## Description + Work on deploying machine learning models to edge devices using frameworks such as Llama.cpp, ONNX, TVM, MLC LLM, and IREE (MLIR). + Collaborate closely with researchers to assist in coding, training and transitioning models from research to production environments. + Integrate AI features into existing products, enriching them with the latest advancements in machine learning. ## Related Videos - [Machine learning in the browser with TensorFlowjs](https://www.wearedevelopers.com/videos/155-machine-learning-in-the-browser-with-tensorflowjs) - [Coffee with Developers - Stephen Jones - NVIDIA](https://www.wearedevelopers.com/videos/1303-coffee-with-developers-stephen-jones-nvidia) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [Developing an AI.SDK](https://www.wearedevelopers.com/videos/198-developing-an-ai-sdk) - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [What Industries Outside of AI Are Hiring The Most AI Experts?](https://www.wearedevelopers.com/magazine/98-what-industries-outside-of-ai-are-hiring-the-most-ai-experts) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence)