AI Devops Infrastructure Engineer/GPU Infrastructure Engineer

Maxonic, Inc.
San Jose, CA, United States
15 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Temporary to permanent
Employment type
Full-time (> 32 hours)
Compensation
$208,000.0 - $257,920.0
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence User Authentication Cloud Engineering Continuous Integration DevOps Programming Tools Firmware Monitoring of Systems Python (Programming Language) Key Management Machine Learning
+19 more
Network Segmentation Performance Tuning Reliability Engineering Ansible Prometheus AI Infrastructure Private Cloud Environment Scripting Grafana Software Troubleshooting Containerization Kubernetes Infrastructure Automation Frameworks Bare Metal Machine Learning Operations Hardware Infrastructure Kibana Terraform Software Version Control

Job description

Maxonic maintains a close and long-term relationship with our direct client. In support of their needs, we are looking for an AI Devops Infrastructure Engineer/GPU Infrastructure Engineer., We are looking for an AI Infrastructure Engineer to help build and operationalize infrastructure supporting AI development, experimentation, training, inference, and AI-enabled developer productivity. This role will focus on building scalable, highly resilient GPU-based infrastructure and enabling engineering teams to access AI resources through reliable, secure, and automated infrastructure services., * Build and operate AI infrastructure supporting development, experimentation, training, and inference.

  • Design and manage infrastructure across GPU and accelerator environments, including NVIDIA and/or AMD.
  • Build scalable infrastructure for GPU capacity planning, utilization, forecasting, workload scheduling, resource allocation, and performance optimization.
  • Develop self-service AI infrastructure capabilities that allow engineering teams to request, provision, manage, extend, and release GPU, CPU, memory, and storage resources.
  • Build and maintain infrastructure across Kubernetes and containerized environments, including compute, networking, storage, and accelerator scheduling.
  • Automate infrastructure provisioning, configuration, scaling, patching, driver/firmware lifecycle management, and decommissioning.
  • Use Infrastructure-as-Code, APIs, automation, and scripting to improve infrastructure reliability and operational efficiency.
  • Establish monitoring, observability, alerting, capacity management, and operational processes for AI infrastructure.
  • Help define and implement best practices for scalable, feasible, and operationally sustainable AI infrastructure.
  • Integrate AI infrastructure with Developer Productivity tooling, CI/CD pipelines, source control, artifact management, build infrastructure, and developer tooling.
  • Partner with engineering, security, IT, and AI/ML teams to make AI resources accessible with minimal operational friction.
  • Troubleshoot complex infrastructure issues and perform root-cause analysis.
  • Help establish secure infrastructure standards covering network segmentation, access controls, authentication, authorization, secrets management, and governance.
  • Evaluate emerging AI infrastructure technologies and improve scalability, reliability, performance, automation, and developer experience.

Requirements

The ideal candidate has hands-on experience building and operating AI/ML infrastructure in real-world environments, with a strong foundation in infrastructure engineering, DevOps/SRE, platform engineering, or similar disciplines., * Strong experience in Infrastructure Engineering, Platform Engineering, DevOps, SRE, Cloud Engineering, or AI/ML Infrastructure.

  • Hands-on experience building and operating AI infrastructure, not just exposure to AI/ML concepts.
  • Strong experience with GPU infrastructure and GPU scalability.
  • Experience with Kubernetes and containerized production environments.
  • Strong understanding of compute, networking, storage, workload scheduling, and resource allocation.
  • Experience with GPU capacity planning, utilization, forecasting, workload scheduling, and performance optimization.
  • Experience building or operating self-service infrastructure/provisioning platforms.
  • Strong experience with Infrastructure-as-Code and automation, such as Terraform, Ansible, Python, APIs, or similar technologies.
  • Experience with CI/CD pipelines and infrastructure operations.
  • Experience with monitoring and observability tools such as Grafana, Prometheus, OpenTelemetry, Kibana, or similar platforms.
  • Experience working across bare-metal, private cloud, and/or public cloud environments.
  • Familiarity with AI/ML infrastructure, inference platforms, and distributed AI workloads.
  • Strong troubleshooting, root-cause analysis, communication, and cross-functional collaboration skills.

About the company

Since 2002 Maxonic has been at the forefront of connecting candidate strengths to client challenges. Our award winning, dedicated team of recruiting professionals are specialized by technology, are great listeners, and will seek to find a position that meets the long-term career needs of our candidates. We take pride in the over 10,000 candidates that we have placed, and the repeat business that we earn from our satisfied clients.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Loading talks and stories from around this role…