Senior Machine Learning Engineer

Connect
London, UK
about 1 month ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Compensation
£70,502.0
Working hours
Regular working hours

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Amazon Web Services Amazon Elastic Compute Cloud Microsoft Azure Cloud Computing Nvidia CUDA Computer Programming Continuous Delivery Continuous Integration Github Python (Programming Language)
+31 more
Load Testing Machine Learning Language Modeling NumPy Performance Tuning Azure Machine Learning Management of Software Versions Speech Recognition WebSocket Google Cloud Chatbots Pytorch Large Language Models Prompt Engineering Multi-Cloud Generative AI HybridCloud Fastapi Containerization Gitlab-ci Scikit Learn Kubernetes Bare Metal Data Analytics Apache Kafka Machine Learning Operations Speech Synthesis Api Design Stream Processing Docker Jenkins

Job description

We are seeking a Senior Machine Learning Engineer to design, deploy, and optimize our next-generation Conversational AI and Data Analytics platforms. You will bridge the gap between AI research and production engineering. Your primary focus will be optimizing and scaling core Speech (ASR, TTS) and Language Model (LLM, SLM) pipelines across hybrid cloud and local edge environments., Pipeline Deployment & Architecture

  • Deploy AI Pipelines: Build production-grade, low-latency pipelines for ASR, TTS, and Small Language Models (SLMs).
  • Hybrid Deployment: Manage deployment topologies across multi-cloud environments and bare-metal local hardware.
  • API Development: Create high-performance, asynchronous REST and WebSocket APIs using FastAPI to serve real-time conversational agents.

MLOps & Infrastructure

  • CI/CD Automation: Design automated machine learning pipelines for model testing, versioning, and continuous deployment.
  • Containerization: Pack applications using Docker or Podman for consistent execution across dev, staging, and production.
  • Multi-Cloud Management: Orchestrate cloud infrastructure across AWS, Azure, and GCP, optimizing for compute efficiency and cost.

Performance Tuning & Optimization

  • GPU Optimization: Maximize hardware utilization for single-GPU and distributed multi-GPU environments.
  • Algorithm Acceleration: Optimize Python code execution using Numba, NumPy, and specialized CUDA libraries.
  • Load Testing: Conduct rigorous load and stress testing to guarantee system stability under high concurrent traffic.

Requirements

Core Programming & Frameworks

  • Language: Mastery of Python and its asynchronous ecosystem.
  • ML Ecosystem: Deep expertise in PyTorch, Scikit-learn, and NumPy.
  • Compilation: Experience accelerating Python code via Numba or Triton.

Conversational AI Experience

  • Speech Technologies: Hands-on experience deploying Automated Speech Recognition (ASR) and Text-to-Speech (TTS) models.
  • Generative AI: Familiarity with optimizing and serving Large Language Models (LLMs) and resource-efficient Small Language Models (SLMs).

Infrastructure & Operations

  • Containers: Advanced knowledge of Docker, Podman, and container orchestration.
  • Cloud Providers: Practical experience managing AI workloads on AWS (EC2, SageMaker), Azure (Azure ML), and GCP (Vertex AI).
  • CI/CD Tools: Experience with GitLab CI, GitHub Actions, Jenkins, or specialized MLOps platforms (e.g., Kubeflow, MLflow)., * Conversational Context: Experience with dialogue management, prompt engineering, and Retrieval-Augmented Generation (RAG).
  • Data Analytics: Familiarity with real-time data streaming (e.g., Kafka) and vector databases (e.g., Pinecone, Milvus, Qdrant).
  • Quantization: Experience with model compression techniques like quantization (INT8/FP4), pruning, and distillation for edge deployment.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.adzuna.co.uk

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:15 min

Open-source community and machine learning frameworks

Gian Marco Iodice Gian Marco Iodice · WWC 2025

2:34 min

Maximizing execution memory effectively via python numpy broadcasting

Jodie Burchell · LIVE

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · WWC 2025

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters · WWC 2023

2:36 min

Exploring high-level Python frameworks for accelerated enterprise artificial intelligence

Paul Graham Paul Graham · LIVE

1:25 min

Replacing NumPy with cuPy for straightforward GPU acceleration

Paul Graham Paul Graham · WWC 2025

Videos

See all

Related articles

See all