> Markdown version of [/jobs/ext/2014126-telecommute-staff-ai-platform-engineer-infrastructure-services](https://www.wearedevelopers.com/jobs/ext/2014126-telecommute-staff-ai-platform-engineer-infrastructure-services). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # TELECOMMUTE Staff AI Platform Engineer, Infrastructure Services - **Company:** · Sentinelone - **Location:** United States (Remote available) - **Experience:** Expert - **Salary:** $156,000.0 - $215,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Code Review, Continuous Integration, DevOps, Programming Tools, Github, OpenID, Package Management Systems, Performance Tuning, AI Infrastructure, Okta, Retrieval-Augmented Generation, Large Language Models, Apigee, Rate Limiting, AI Platforms, Kubernetes, Low Latency, Github Enterprise, Machine Learning Operations, Nim (Programming Language), Api Design, Api Gateway, Terraform, Software Version Control, Jenkins, Artifactory - **Published:** August 10, 2026 - **Apply:** https://www.dice.com/job-detail/3be226a5-4f23-4a6c-83d3-2a7bfbc3ecb3 ## About the Role * 8 or more years of experience in platform, infrastructure, or DevOps engineering, with a track record of owning systems end-to-end in production. * Hands-on experience with API gateway technologies (Kong, Envoy, Apigee, or similar); direct experience with AI/LLM gateway patterns (rate limiting, semantic caching, prompt/response observability) is a strong plus. * Strong Kubernetes and GitOps experience (ArgoCD or comparable), and comfort operating across multiple environments (dev, gov, prod). * Solid CI/CD background: Jenkins pipeline design and administration, build infrastructure, and runner/agent fleet management (GitHub Actions runners or equivalent). * Experience with artifact and package management systems (Artifactory, Xray, or similar) and source control platform administration (GitHub Enterprise). * Working knowledge of infrastructure-as-code (Terraform) and cloud platforms (AWS/EKS). * Experience deploying and operating self-hosted LLM inference stacks (vLLM, NVIDIA Triton/NIM, TGI, Ollama, or similar) and GPU-backed infrastructure, including Kubernetes GPU scheduling and autoscaling. * Familiarity with LLMOps practices: model versioning, evaluation harnesses, and usage/cost observability across API-based and self-hosted models. * Track record of setting technical direction, driving cross-team initiatives, and mentoring other engineers; this role has significant scope and minimal day-to-day oversight. * Clear, proactive communicator who can explain infrastructure trade-offs to both engineers and non-technical stakeholders. * Experience operating LLM/AI-assisted developer tooling at scale (Claude Code, Copilot, or similar) inside an enterprise is preferred. * Familiarity with Okta/OIDC and enterprise auth patterns for internal platforms is preferred. * Experience with engineering productivity metrics tooling (LinearB or similar) and AI-based code review tooling (Qodo or similar) is preferred. * Experience with vector databases and RAG pipelines (e.g. Milvus, Pinecone, pgvector, or similar) in a production setting is preferred. * Exposure to model fine-tuning or lightweight training pipelines (LoRA/QLoRA or similar) for domain-specific model adaptation is preferred. ## Description As a Staff AI Platform Engineer, Infrastructure Services, you will be tasked with taking ownership of our AI Gateway infrastructure (built on Kong AI Gateway), the system that authenticates, routes, rate-limits, and monitors AI coding assistant traffic org-wide, while also being fluent enough across our broader platform stack to design solutions that span the two. This is a high-autonomy, high-scope role: you will set technical direction for AI infrastructure, drive incident response and reliability work, and partner closely with the engineers who own our CI/CD, GitOps, and artifact systems rather than working in isolation from them. What Will You Do? Primary responsibilities include: * Work on the AI Gateway platform: architect, harden, and scale our Kong AI Gateway deployment (Konnect Hybrid on KCP/EKS), including auth (Okta/OIDC), consumer tiers and budgets, rate limiting, semantic caching, and observability. * Lead reliability and incident response: drive root-cause analysis and remediation for gateway issues (timeouts, latency, capacity, failover) and build the monitoring/alerting needed to catch them before users do. * Design across the platform, not just the gateway: work fluently with our CI/CD (Jenkins, JPAAS), GitOps and Kubernetes deployment tooling (ArgoCD across dev/gov/prod), artifact management (Artifactory/Xray), GitHub Enterprise administration, and GitHub Actions runner fleet, so that AI infrastructure decisions account for how the rest of the platform actually works. * Evaluate and roll out AI developer tooling: run structured pilots and adoption efforts for tools like AI-assisted PR review (Qodo) and engineering metrics platforms (LinearB), and make clear build-vs-buy recommendations. * Set technical direction and mentor: define architecture and standards for AI infrastructure, review designs across the team, and raise the bar for other engineers working in this space. * Partner cross-functionally: work directly with security, DevEx, and product engineering teams consuming the gateway to translate their needs into platform capabilities. * Host and serve local models: stand up and operate self-hosted/open-weight model serving infrastructure (e.g. vLLM, NVIDIA Triton/NIM, TGI, Ollama) for workloads where routing to an external provider isn't the right fit, including GPU capacity planning, autoscaling, and cost/performance tuning. * Support the broader model lifecycle: help build LLMOps practices such as model versioning, evaluation, and safe rollout, plus supporting infrastructure for retrieval-augmented generation (vector stores, embedding pipelines) as use cases mature. * Track usage and cost: build observability into token usage, latency, and spend across both API-based and self-hosted models so the business can see what AI infrastructure actually costs. ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [DevOps for AI: running LLMs in production with Kubernetes and KubeFlow](https://www.wearedevelopers.com/videos/1222-devops-for-ai-running-llms-in-production-with-kubernetes-and-kubeflow) - [Agentic employees in world's most downloaded FinTech app](https://www.wearedevelopers.com/videos/100123-agentic-employees-in-world-s-most-downloaded-fintech-app) - [Bringing AI Model Testing and Prompt Management to Your Codebase with GitHub Models](https://www.wearedevelopers.com/videos/1536-bringing-ai-model-testing-and-prompt-management-to-your-codebase-with-github-models) ## Related Articles - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere)