> Markdown version of [/jobs/ext/3387006-ai-devops-infrastructure-engineer-gpu-infrastructure-engineer](https://www.wearedevelopers.com/jobs/ext/3387006-ai-devops-infrastructure-engineer-gpu-infrastructure-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Devops Infrastructure Engineer/GPU Infrastructure Engineer - **Company:** Maxonic, Inc. - **Location:** San Jose, CA, United States - **Salary:** $208,000.0 - $257,920.0 - **Contract:** Temporary to permanent - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, User Authentication, Cloud Engineering, Continuous Integration, DevOps, Programming Tools, Firmware, Monitoring of Systems, Python (Programming Language), Key Management, Machine Learning, Network Segmentation, Performance Tuning, Reliability Engineering, Ansible, Prometheus, AI Infrastructure, Private Cloud Environment, Scripting, Grafana, Software Troubleshooting, Containerization, Kubernetes, Infrastructure Automation Frameworks, Bare Metal, Machine Learning Operations, Hardware Infrastructure, Kibana, Terraform, Software Version Control - **Published:** September 22, 2026 - **Apply:** https://www.dice.com/job-detail/3f597fb4-ded2-4875-b679-59a6b81de995 ## About the Role The ideal candidate has hands-on experience building and operating AI/ML infrastructure in real-world environments, with a strong foundation in infrastructure engineering, DevOps/SRE, platform engineering, or similar disciplines., * Strong experience in Infrastructure Engineering, Platform Engineering, DevOps, SRE, Cloud Engineering, or AI/ML Infrastructure. * Hands-on experience building and operating AI infrastructure, not just exposure to AI/ML concepts. * Strong experience with GPU infrastructure and GPU scalability. * Experience with Kubernetes and containerized production environments. * Strong understanding of compute, networking, storage, workload scheduling, and resource allocation. * Experience with GPU capacity planning, utilization, forecasting, workload scheduling, and performance optimization. * Experience building or operating self-service infrastructure/provisioning platforms. * Strong experience with Infrastructure-as-Code and automation, such as Terraform, Ansible, Python, APIs, or similar technologies. * Experience with CI/CD pipelines and infrastructure operations. * Experience with monitoring and observability tools such as Grafana, Prometheus, OpenTelemetry, Kibana, or similar platforms. * Experience working across bare-metal, private cloud, and/or public cloud environments. * Familiarity with AI/ML infrastructure, inference platforms, and distributed AI workloads. * Strong troubleshooting, root-cause analysis, communication, and cross-functional collaboration skills. ## Description Maxonic maintains a close and long-term relationship with our direct client. In support of their needs, we are looking for an AI Devops Infrastructure Engineer/GPU Infrastructure Engineer., We are looking for an AI Infrastructure Engineer to help build and operationalize infrastructure supporting AI development, experimentation, training, inference, and AI-enabled developer productivity. This role will focus on building scalable, highly resilient GPU-based infrastructure and enabling engineering teams to access AI resources through reliable, secure, and automated infrastructure services., * Build and operate AI infrastructure supporting development, experimentation, training, and inference. * Design and manage infrastructure across GPU and accelerator environments, including NVIDIA and/or AMD. * Build scalable infrastructure for GPU capacity planning, utilization, forecasting, workload scheduling, resource allocation, and performance optimization. * Develop self-service AI infrastructure capabilities that allow engineering teams to request, provision, manage, extend, and release GPU, CPU, memory, and storage resources. * Build and maintain infrastructure across Kubernetes and containerized environments, including compute, networking, storage, and accelerator scheduling. * Automate infrastructure provisioning, configuration, scaling, patching, driver/firmware lifecycle management, and decommissioning. * Use Infrastructure-as-Code, APIs, automation, and scripting to improve infrastructure reliability and operational efficiency. * Establish monitoring, observability, alerting, capacity management, and operational processes for AI infrastructure. * Help define and implement best practices for scalable, feasible, and operationally sustainable AI infrastructure. * Integrate AI infrastructure with Developer Productivity tooling, CI/CD pipelines, source control, artifact management, build infrastructure, and developer tooling. * Partner with engineering, security, IT, and AI/ML teams to make AI resources accessible with minimal operational friction. * Troubleshoot complex infrastructure issues and perform root-cause analysis. * Help establish secure infrastructure standards covering network segmentation, access controls, authentication, authorization, secrets management, and governance. * Evaluate emerging AI infrastructure technologies and improve scalability, reliability, performance, automation, and developer experience.