Infrastructure Engineer

ITcaps LLC
St. Louis, United States
10 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Ubuntu (Operating System) Linux Prometheus Azure Machine Learning Runbook Ceph (Software) Graphics Processing Unit (GPU) Cloud Platform System Grafana Software Troubleshooting Kubernetes
+4 more
Infrastructure Automation Frameworks Machine Learning Operations Hardware Infrastructure Terraform

Job description

  • Strong Kubernetes experience
  • Linux/Ubuntu expertise
  • NVIDIA GPUs and NVIDIA stack
  • Terraform
  • GPU-enabled Kubernetes infrastructure

Responsibilities:

  • Design and manage Kubernetes clusters
  • Build and manage GPU-enabled infrastructure
  • Deploy and manage Longhorn storage
  • Automate infrastructure using Terraform
  • Monitor platforms using Prometheus and Grafana
  • Troubleshoot and optimize AI/ML infrastructure, * Conduct KT sessions on Kubernetes, GPU infrastructure, NVIDIA stack, and storage
  • Create technical documentation, runbooks, and training materials
  • Conduct hands-on workshops for client teams
  • Advise on cloud-native infrastructure, AI/ML operations, reliability, performance, and cost optimization
  • Ensure smooth production handoff and operational readiness

Requirements

  • Longhorn or Ceph
  • Canonical MAAS, Juju, or Charmed Kubernetes
  • Kubeflow or other AI/ML platforms
  • CKA or NVIDIA certifications, * Strong troubleshooting and problem-solving
  • Excellent communication and documentation
  • Ability to collaborate with ML engineers and data scientists
  • Strong focus on automation and platform scalability

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

2:50 min

Introduction and the value of runbooks

Hila Fish · World Congress 2023

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

1:24 min

Comprehensive AI infrastructure stacks at the Linux Foundation

Matt White Matt White · World Congress 2025

1:04 min

Visualizing Keycloak performance via standard Grafana troubleshooting dashboards

Alexander Schwartz Alexander Schwartz · World Congress 2025

Videos

See all

Related articles

See all