> Markdown version of [/jobs/ext/258142-devops-engineer-ai-infrastructure-platforms](https://www.wearedevelopers.com/jobs/ext/258142-devops-engineer-ai-infrastructure-platforms). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # DevOps Engineer - AI Infrastructure & Platforms - **Company:** Insight Global - **Location:** Palo Alto, CA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Kubernetes Security, Artificial Intelligence, Amazon Web Services, Automation of Tests, Microsoft Azure, Cloud Computing, Cloud Engineering, Continuous Integration, DevOps, Monitoring of Systems, Identity and Access Management, Python (Programming Language), Node.Js, Open Web Application Security, Shell Script, Web Applications, AI Infrastructure, Pulumi, Data Processing, ReactJS, Delivery Pipeline, Large Language Models, Model Validation, Infrastructure as Code (IaC), Cloudformation, Containerization, AI Platforms, Kubernetes, Machine Learning Operations, Terraform, Data Pipelines - **Published:** May 17, 2026 - **Apply:** https://www.juju.com/job/00000000g0ctqg ## About the Role * 6+ years of DevOps/SRE experience in a cloud-native environment. * Expert-level Kubernetes (K8s) knowledge, including cluster security, networking, and scaling. * Strong proficiency in Python (for automation scripts and data pipeline support) and Shell scripting. * Hands-on experience with Cloud Providers: Deep expertise in AWS, GCP, or Azure. * IaC Mastery: Proven experience with Terraform, Pulumi, or similar tools. * Security Mindset: Experience securing applications in Kubernetes and familiarity with container security scanning. * Prior experience supporting ML/AI teams or managing GPU-accelerated workloads. * Experience with MLOps tools (e.g., Kubeflow, MLflow, or Weights & Biases). * Familiarity with Vector Databases or high-scale data processing engines. * Background in automating complex simulation environments or sandboxes. ## Description We are seeking a Senior DevOps Engineer to join our specialized AI Engineering and Research team. This team is responsible for building Everse (an evaluation and simulation platform for AI agents) and advanced LLM data pipelines. Your role will focus on architecting the underlying infrastructure that allows our researchers and engineers to deploy, scale, and monitor complex AI models and web applications securely. You will bridge the gap between AI research and production-grade stability, ensuring our Kubernetes clusters and CI/CD pipelines are optimized for high-performance AI workloads., * Infrastructure as Code (IaC): Design, build, and maintain scalable cloud infrastructure using Terraform or CloudFormation. * Kubernetes Orchestration: Manage and optimize secure Kubernetes clusters, specifically for hosting data-heavy React/Node.js applications and Python-based AI services. * CI/CD Pipeline Development: Build and automate robust deployment pipelines to ensure rapid, high-frequency releases for the Everse platform. * MLOps Support: Collaborate with AI scientists to streamline the deployment of LLM and RLHF workflows, managing the infrastructure required for model evaluation and simulation. * Security & Compliance: Implement security best practices (OWASP, IAM roles) to ensure data privacy within our annotation and video surveillance tools. * Monitoring & Observability: Establish deep visibility into system performance and cost-tracking for cloud resources (AWS/GCP/Azure). ## Related Videos - [Watch Tests Go Brrrr! : Getting Started with Cypress in ReactJS](https://www.wearedevelopers.com/videos/282-watch-tests-go-brrrr-getting-started-with-cypress-in-reactjs) - [Unleashing Potential Across Teams: The Power of Infrastructure as Code](https://www.wearedevelopers.com/videos/930-unleashing-potential-across-teams-the-power-of-infrastructure-as-code) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [DevOps for AI: running LLMs in production with Kubernetes and KubeFlow](https://www.wearedevelopers.com/videos/1222-devops-for-ai-running-llms-in-production-with-kubernetes-and-kubeflow) - [Terraform for Developers](https://www.wearedevelopers.com/videos/3-terraform-for-developers) ## Related Articles - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers)