> Markdown version of [/jobs/ext/340708-machine-learning-engineer](https://www.wearedevelopers.com/jobs/ext/340708-machine-learning-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Machine Learning Engineer - **Company:** Signapse - **Location:** Stockport, UK - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Profiling, Software Debugging, Python (Programming Language), Machine Learning, Data Streaming, Management of Software Versions, Graphics Processing Unit (GPU), Deep Learning, Backend, Kubernetes, Infrastructure Automation Frameworks, ONNX (Open Neural Network Exchange) Format, Performance Monitor, Hardware Acceleration, Machine Learning Operations, TensorRT, Hardware Infrastructure, Docker - **Published:** June 13, 2026 - **Apply:** https://www.apply4u.co.uk/jobs/x/38946154/ ## About the Role What we're looking forEssential3+ years of experience in ML systems engineering, ML infrastructure, or backend systems.Strong Python skills (Rust is a bonus)Experience working with production ML modelsStrong debugging, profiling, and performance analysis skillsA genuine interest in building latency-critical, high-throughput systemsDesirableExperience with TensorRT, ONNX, Triton, TorchServe, or similar inference tools.Familiarity with GPU architecture and performance optimisationExperience with video, graphics, or real-time streaming systems (HLS, RTMP,SRT)Experience with Kubernetes, Docker, and ML workloads at scale.Familiarity with AWS, including Sage Maker. ## Description The ChallengeOur generative AI models produce sign language video in real time, delivered over GPU infrastructure across cloud and on-prem environments to a global audience. The engineering problem is genuinely hard: keep latency low, maximise GPU utilisation, and build infrastructure that scales to hundreds of simultaneous streams. You'll work across the full ML inference stack from model optimisation to deployment infrastructure and own this challenge. What you'll work onML inference optimisationProfile and optimise deep learning models used for sign language video generationReduce inference latency using quantisation, pruning, mixed precision, and kernel optimisationImprove GPU utilisation and throughput across inference pipelinesWork closely with ML researchers to ensure models are production-readyML infrastructure & deploymentBuild and maintain scalable model serving systems on GPU clustersDesign autoscaling infrastructure to meet real-time SLAsContribute to model deployment pipelines, versioning, and rollback strategiesPerformance engineeringDevelop benchmarking frameworks for tracking inference performanceIdentify bottlenecks across the ML pipeline and eliminate latency hotspotsImplement performance monitoring and alerting for production systemsEvaluate new hardware accelerators and inference run times. Scaling globallyWork with the research team to expand sign languages and digital signersArchitect systems that allow rapid onboarding of new languagesBuild low-latency infrastructure that scales to hundreds of concurrent streams ## Related Videos - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Introduction to Azure Machine Learning](https://www.wearedevelopers.com/videos/368-introduction-to-azure-machine-learning) - [Developing the Backend with Stefan Lingler, CTO at Shpock](https://www.wearedevelopers.com/videos/100360-developing-the-backend-with-stefan-lingler-cto-at-shpock) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [Nest.js - TypeScript in the backend can also be clean](https://www.wearedevelopers.com/videos/1033-nest-js-typescript-in-the-backend-can-also-be-clean) ## Related Articles - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models)