AI Infrastructure Operations Bootcamp Instructor

WeCloudData
UK
2 months ago

Role details

Contract type
Permanent contract
Employment type
Part-time (≤ 32 hours)
Experience level
Experienced
Experience required
2 years minimum
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Amazon Web Services Microsoft Azure Cloud Computing Cloud Engineering Nvidia CUDA Data Centers Linux DevOps Monitoring of Systems Image Management Python (Programming Language)
+22 more
Laboratory Information Management Systems Linux System Administration Machine Learning Networking Basics Reliability Engineering Prometheus Shell Script Software Deployment AI Infrastructure Data Logging Scripting Graphics Processing Unit (GPU) Google Cloud Autoscaling Large Language Models Grafana Containerization AI Platforms Kubernetes Machine Learning Operations Nim (Programming Language) Docker

Job description

WeCloudData is seeking an experienced AI Infrastructure Operations Instructor to deliver a hands-on bootcamp focused on GPU-enabled AI systems, Kubernetes operations, AI deployment, observability, and infrastructure monitoring. The instructor will train students to deploy, operate, monitor, and troubleshoot AI workloads in modern cloud and data center environments. This role focuses on the infrastructure and operational side of AI systems rather than AI model development or research. The ideal candidate has experience operating production AI platforms, deploying containerized applications, managing Kubernetes environments, and working with GPU-enabled infrastructure. The instructor should be comfortable teaching beginners and helping students transition into AI Infrastructure Operations, MLOps, and AI Platform Engineering careers., Instruction & Delivery

  • Deliver instructor-led lectures and workshops
  • Conduct hands-on labs and troubleshooting exercises
  • Mentor students throughout the bootcamp
  • Support students in completing capstone projects
  • Evaluate assignments and provide technical feedback

Curriculum Coverage

  • Teach topics including:
  • AI infrastructure fundamentals
  • Training vs inference workloads
  • GPU fundamentals and operations
  • Linux administration
  • Python scripting
  • Containerization using Docker
  • AI model deployment and serving
  • Kubernetes operations
  • GPU workload scheduling
  • Monitoring and observability
  • AI platform operations
  • AI infrastructure troubleshooting
  • AI data center fundamentals

  • The instructor should be capable of guiding students through all major topics outlined in the bootcamp curriculum. Lab Management

  • Prepare cloud-based lab environments
  • Configure Kubernetes clusters
  • Manage GPU-enabled infrastructure
  • Support Docker and container deployment labs
  • Create troubleshooting scenarios and exercises
  • Maintain capstone project environments

Requirements

Do you have experience in Python?, Do you have a Master’s degree?, 5+ years of experience in one or more of the following:

  • Platform Engineering
  • Cloud Engineering
  • Site Reliability Engineering (SRE)
  • MLOps
  • AI Platform Operations
  • Infrastructure Engineering
  • DevOps Engineering

Preferred

  • 2+ years supporting AI or machine learning workloads in production

Technical Expertise

  • Linux
  • Strong hands-on experience with:
  • Linux administration
  • Shell scripting
  • Process management
  • Networking fundamentals
  • System troubleshooting

  • Containers
  • Strong experience with:
  • Docker
  • Container image management
  • Containerized application deployment

  • Kubernetes
  • Practical experience with:
  • Pods
  • Deployments
  • Services
  • Ingress
  • Autoscaling
  • Helm
  • Monitoring Kubernetes workloads

  • Cloud Platforms
  • Experience with one or more:
  • AWS
  • Azure
  • Google Cloud
  • Alibaba Cloud

  • Monitoring & Observability
  • Experience with:
  • Prometheus
  • Grafana
  • Logging platforms
  • Alerting systems
  • Infrastructure monitoring

AI Infrastructure Knowledge

  • Must understand:
  • Training vs inference workloads
  • AI deployment architectures
  • Model serving concepts
  • GPU utilization concepts
  • AI workload bottlenecks
  • Throughput and latency metrics
  • AI platform operations

  • The instructor does not need to be an AI researcher but should be comfortable deploying and operating AI workloads.

Preferred Qualifications GPU & AI Infrastructure

  • Experience with:
  • NVIDIA GPUs
  • CUDA ecosystem
  • GPU monitoring tools
  • GPU scheduling concepts
  • Multi-GPU systems

  • Bonus
  • NVIDIA AI Enterprise
  • NVIDIA NIM
  • Triton Inference Server
  • vLLM
  • Ray Serve

MLOps & AI Platform Experience

  • Experience with:
  • MLflow
  • Kubeflow
  • Model serving platforms
  • Vector databases
  • RAG deployment architectures
  • LLM inference systems

  • These align strongly with the capstone projects proposed in the curriculum.

Preferred Certifications Strongly Preferred

  • Kubernetes Administrator (CKA)
  • Kubernetes Application Developer (CKAD)

About the company

WeCloudData is a premier learning academy dedicated to offering top-quality data and AI training to our students and corporate clients. Recognized as a leading school in the data and AI category, we have expanded our offerings to include consultancy and career services, supporting our students beyond the classroom. Over the next decade, our goal is to influence millions of learners, driving positive changes in the future of work through our commitment to excellent learning experiences, technological innovation, and fostering work-integrated learning environments. We pride ourselves on setting industry standards and being at the forefront of data and AI education.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:24 min

Comprehensive AI infrastructure stacks at the Linux Foundation

Matt White Matt White · WWC 2025

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · WWC 2025

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · WWC 2025

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

2:48 min

Exploring generative AI learning paths and hands-on labs

Asrar Asrar · WWC 2024

2:39 min

Experiencing core Linux capabilities for DevOps administration

Michael Cade · LIVE

Videos

See all

Related articles

See all