> Markdown version of [/jobs/ext/2010601-ai-platform-engineer-infrastructure-services](https://www.wearedevelopers.com/jobs/ext/2010601-ai-platform-engineer-infrastructure-services). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Platform Engineer, Infrastructure Services - **Company:** · Sentinelone - **Location:** United States (Remote available) - **Experience:** Expert - **Salary:** $132,000.0 - $182,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Code Review, Continuous Integration, DevOps, Programming Tools, Github, OpenID, Package Management Systems, Performance Tuning, AI Infrastructure, Okta, Retrieval-Augmented Generation, Large Language Models, Apigee, Rate Limiting, Kubernetes, Low Latency, Github Enterprise, Machine Learning Operations, Nim (Programming Language), Api Design, Api Gateway, Terraform, Software Version Control, Jenkins, Artifactory - **Published:** August 10, 2026 - **Apply:** https://www.dice.com/job-detail/14785f82-7bae-437c-9e57-e774f236fc79 ## About the Role * 5 or more years of experience in platform, infrastructure, or DevOps engineering, with a track record of owning systems end-to-end in production. * Hands-on experience with API gateway technologies (Kong, Envoy, Apigee, or similar); direct experience with AI/LLM gateway patterns (rate limiting, semantic caching, prompt/response observability) is a strong plus. * Strong Kubernetes and GitOps experience (ArgoCD or comparable), and comfort operating across multiple environments (dev, gov, prod). * Solid CI/CD background: Jenkins pipeline design and administration, build infrastructure, and runner/agent fleet management (GitHub Actions runners or equivalent). * Experience with artifact and package management systems (Artifactory, Xray, or similar) and source control platform administration (GitHub Enterprise). * Working knowledge of infrastructure-as-code (Terraform) and cloud platforms (AWS/EKS). * Experience deploying and operating self-hosted LLM inference stacks (vLLM, NVIDIA Triton/NIM, TGI, Ollama, or similar) and GPU-backed infrastructure, including Kubernetes GPU scheduling and autoscaling. * Familiarity with LLMOps practices: model versioning, evaluation harnesses, and usage/cost observability across API-based and self-hosted models. * Track record of setting technical direction, driving cross-team initiatives, and mentoring other engineers; this role has significant scope and minimal day-to-day oversight. * Clear, proactive communicator who can explain infrastructure trade-offs to both engineers and non-technical stakeholders. * Experience operating LLM/AI-assisted developer tooling at scale (Claude Code, Copilot, or similar) inside an enterprise is preferred. * Familiarity with Okta/OIDC and enterprise auth patterns for internal platforms is preferred. * Experience with engineering productivity metrics tooling (LinearB or similar) and AI-based code review tooling (Qodo or similar) is preferred. * Experience with vector databases and RAG pipelines (e.g. Milvus, Pinecone, pgvector, or similar) in a production setting is preferred. * Exposure to model fine-tuning or lightweight training pipelines (LoRA/QLoRA or similar) for domain-specific model adaptation is preferred. ## Description As a Senior AI Platform Engineer, Infrastructure Services, you will be tasked with taking ownership of our AI Gateway infrastructure (built on Kong AI Gateway), the system that authenticates, routes, rate-limits, and monitors AI coding assistant traffic org-wide, while also being fluent enough across our broader platform stack to design solutions that span the two. This is a high-autonomy, high-scope role: you will set technical direction for AI infrastructure, drive incident response and reliability work, and partner closely with the engineers who own our CI/CD, GitOps, and artifact systems rather than working in isolation from them. What Will You Do? Primary responsibilities include: * Work on the AI Gateway platform: architect, harden, and scale our Kong AI Gateway deployment (Konnect Hybrid on KCP/EKS), including auth (Okta/OIDC), consumer tiers and budgets, rate limiting, semantic caching, and observability. * Lead reliability and incident response: drive root-cause analysis and remediation for gateway issues (timeouts, latency, capacity, failover) and build the monitoring/alerting needed to catch them before users do. * Design across the platform, not just the gateway: work fluently with our CI/CD (Jenkins, JPAAS), GitOps and Kubernetes deployment tooling (ArgoCD across dev/gov/prod), artifact management (Artifactory/Xray), GitHub Enterprise administration, and GitHub Actions runner fleet, so that AI infrastructure decisions account for how the rest of the platform actually works. * Evaluate and roll out AI developer tooling: run structured pilots and adoption efforts for tools like AI-assisted PR review (Qodo) and engineering metrics platforms (LinearB), and make clear build-vs-buy recommendations. * Set technical direction and mentor: define architecture and standards for AI infrastructure, review designs across the team, and raise the bar for other engineers working in this space. * Partner cross-functionally: work directly with security, DevEx, and product engineering teams consuming the gateway to translate their needs into platform capabilities. * Host and serve local models: stand up and operate self-hosted/open-weight model serving infrastructure (e.g. vLLM, NVIDIA Triton/NIM, TGI, Ollama) for workloads where routing to an external provider isn't the right fit, including GPU capacity planning, autoscaling, and cost/performance tuning. * Support the broader model lifecycle: help build LLMOps practices such as model versioning, evaluation, and safe rollout, plus supporting infrastructure for retrieval-augmented generation (vector stores, embedding pipelines) as use cases mature. * Track usage and cost: build observability into token usage, latency, and spend across both API-based and self-hosted models so the business can see what AI infrastructure actually costs. ## Related Videos - [DevOps for AI: running LLMs in production with Kubernetes and KubeFlow](https://www.wearedevelopers.com/videos/1222-devops-for-ai-running-llms-in-production-with-kubernetes-and-kubeflow) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Bringing AI Model Testing and Prompt Management to Your Codebase with GitHub Models](https://www.wearedevelopers.com/videos/1536-bringing-ai-model-testing-and-prompt-management-to-your-codebase-with-github-models) - [DevOps Maturity Check – a way to balance autonomy and alignment](https://www.wearedevelopers.com/videos/58-devops-maturity-check-a-way-to-balance-autonomy-and-alignment) ## Related Articles - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift)