Platform Engineer

Virtual Networx
St. Louis, MO, United States
1 day ago

Role details

Contract type
Temporary contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Ubuntu (Operating System) Performance Tuning Prometheus Azure Machine Learning Ceph (Software) Graphics Processing Unit (GPU) Data Storage Management Cloud Platform System Grafana Kubernetes Infrastructure Automation Frameworks
+3 more
Machine Learning Operations Hardware Infrastructure Terraform

Job description

  • Design and manage Kubernetes clusters
  • Build GPU-enabled infrastructure
  • Deploy Longhorn storage
  • Automate infrastructure using Terraform
  • Monitor systems using Prometheus and Grafana
  • Knowledge Transfer & Client Enablement
  • Provide structured knowledge transfer (KT) sessions to client teams on all core platform components, including:
  • Kubernetes architecture, operations, and troubleshooting
  • GPU infrastructure (NVIDIA stack, scheduling, resource optimization)
  • Longhorn storage management and performance tuning
  • Canonical ecosystem tools (MAAS, Juju, Charmed Kubernetes)

  • Develop and deliver technical documentation, runbooks, and training materials to support ongoing operations
  • Conduct hands-on workshops and guided sessions to enable client teams to independently manage and scale the platform
  • Act as a technical advisor, helping client stakeholders understand best practices in:

  • Cloud-native infrastructure o AI/ML platform operations o Reliability, performance, and cost optimization

  • Ensure smooth handoff of production systems with full operational readiness and support knowledge

Requirements

  • 3 8+ years’ experience
  • Strong Kubernetes knowledge
  • Experience with GPUs and NVIDIA stack
  • Linux (Ubuntu) expertise
  • Experience with Terraform

Preferred Qualifications

  • Longhorn or Ceph experience
  • Canonical ecosystem (MAAS, Juju)
  • AI/ML tools like Kubeflow
  • Certifications (CKA, NVIDIA)

Soft Skills

  • Strong problem-solving and troubleshooting mindset
  • Ability to collaborate with cross-functional teams (ML engineers, data scientists)
  • Clear communication and documentation skills
  • Passion for automation and platform scalability

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com

Good distractions

Talks and stories from around this role β€” technically off-topic, practically not.

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet Β· LIVE

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo Β· LIVE

1:05 min

Measuring system availability utilizing Prometheus and straightforward PromQL

Alexander Schwartz Alexander Schwartz Β· World Congress 2025

1:34 min

Essential commands for running and testing Terraform configurations

Hennie Francis Β· LIVE

13:07 min

Configuring application observability with Micrometer and Prometheus

Aleksandr Kalikov Β· LIVE

4:47 min

Automating frontend performance metrics with Google Lighthouse

Miki Lombardi Β· JS Congress

Videos

See all

Related articles

See all