Senior Deep Learning Engineer

Nvidia
UK
18 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
£221,250.0
Working hours
Regular working hours

Tech stack

Artificial Intelligence Nvidia CUDA Computer Programming Python (Programming Language) Performance Tuning Software Deployment Pytorch Large Language Models Deep Learning Optimization Algorithms Machine Learning Operations TensorRT
+1 more
Docker

Job description

We are looking for a Senior Deep Learning Engineer to help bring Cosmos World Foundation Models from research into efficient, production-grade systems. You’ll focus on optimizing and deploying models for high-performance inference on diverse GPU platforms. This role sits at the intersection of deep learning, systems, and GPU optimization - working closely with research scientists, software engineers, and hardware experts.

NVIDIA Cosmos is a platform purpose-built for physical AI, featuring powerful generative models. Developers use Cosmos to accelerate physical AI development for autonomous vehicles (AVs), robots, and video analytics AI agents by simulating and reasoning about the physical world.

What you’ll be doing:

  • Improve inference speed for Cosmos WFMs on GPU platforms.
  • Effectively carry out the production deployment of Cosmos WFMs.
  • Profile and analyze deep learning workloads to identify and remove bottlenecks.

Requirements

  • 5+ years of experience.
  • MSc or PhD in CS, EE, or CSEE or equivalent experience.
  • Strong background in Deep Learning.
  • Strong programming skills in Python and PyTorch.
  • Experience with inference optimization techniques (such as quantization) and inference optimization frameworks, one of: TensorRT, TensorRT-LLM, vLLM, SGLang.

Ways to stand out from the crowd:

  • Familiarity with deploying Deep Learning models in production settings (e.g., Docker, Triton Inference Server).
  • CUDA programming experience.
  • Familiarity with diffusion models.
  • Proven experience in analyzing, modeling, and tuning the performance of GPU workloads, both inference and training.

Benefits & conditions

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. For Poland: The base salary range is 221,250 PLN - 383,500 PLN for Level 3, and 292,500 PLN - 507,000 PLN for Level 4.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.totaljobs.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:33 min

Architecting CUDA and the AI software stack

Michael Kagan Michael Kagan +1 · WWC Europe 2026

4:52 min

Essential phases in building and refining language models

Anshul Jindal Anshul Jindal +1 · WWC 2025

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · WWC 2023

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · WWC 2025

2:08 min

History and scale of NVIDIA GPU computing

Paul Graham Paul Graham · LIVE

2:32 min

Core libraries driving inference engines and multi-GPU networking

Adolf Hohl Adolf Hohl · WWC 2024

Videos

See all

Related articles

See all