Platform Engineer - AI Infrastructure

Stack Infrastructure
Denver, CO, United States
27 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
2 years minimum
Compensation
$128,260.0 - $146,018.0
Working hours
Regular working hours
Job source

Tech stack

A/B Testing Application Programming Interfaces (APIs) Artificial Intelligence Automation of Tests Microsoft Azure Bash Shell Cloud Computing Continuous Integration Data Infrastructure DevOps Domain Name System (DNS) Github
+29 more
Python (Programming Language) Role-Based Access Control Azure Data Lake Search Technologies Software Deployment Management of Software Versions AI Infrastructure Cloud Platform System Microsoft Power Automate Autoscaling Delivery Pipeline Large Language Models Kubernetes Helm Charts Microsoft InTune Fastapi Containerization AI Platforms Kubernetes Infrastructure Automation Frameworks Deployment Automation Bicep Cosmos DB Machine Learning Operations Terraform Data Pipelines Serverless Computing Docker Key Vault Databricks

Job description

Platform Engineer - AI Infrastructure owns the cloud platform, DevOps pipelines, automation runtime environments, and operational infrastructure that power all of STACK’s AI, automation, and data initiatives. This is a hands-on leadership role-responsible for ensuring that every intelligent agent, automation workflow, RAG platform, and data pipeline moves from prototype to production rapidly, runs reliably, and scales cost-effectively. The scope spans Azure infrastructure provisioning using Terraform and Bicep, CI/CD pipeline engineering with Azure DevOps and GitHub Actions, container orchestration on AKS and Azure Container Apps, model serving and vector search infrastructure, automation runtime hosting, security hardening, and FinOps cost management. This lead also owns the deployment infrastructure for agentic and hybrid model workloads-including LLM/SLM serving endpoints, embedding compute, GPU/inference scaling, and multi-model routing. The ideal candidate is equally comfortable writing Terraform modules and reviewing architecture diagrams, with a relentless focus on deployment velocity, reliability, cost optimization, and security.

Azure Infrastructure & Platform Engineering

  • Design, deploy, and manage Azure infrastructure across dual EA subscriptions (Dev/Non-Prod and Production) including Databricks workspaces, AI Search clusters, Cosmos DB instances, ADLS Gen2, Azure OpenAI Service endpoints, and Azure Functions.
  • Implement Infrastructure-as-Code using Terraform, Bicep, or ARM templates with modular, version-controlled patterns enabling new workloads to deploy within hours.
  • Configure Azure networking (VNets, Private Endpoints, NSGs, Private DNS) for secure, globally distributed platform environments across AMER, EMEA, and APAC.
  • Build container-based deployment patterns (Azure Container Apps, AKS) for API serving, agent hosting, model inference, and automation execution.
  • Provision and manage LLM/SLM serving infrastructure: Azure OpenAI deployments, model endpoints, token-based scaling, and multi-region failover.

CI/CD, MLOps & Automation Runtime

  • Design end-to-end CI/CD pipelines (Azure DevOps, GitHub Actions) for application deployment, model promotion, data pipeline orchestration, and automated testing with blue/green and canary patterns.
  • Build MLOps pipelines for model registration, versioning, A/B testing, canary deployment, and automated rollback of LLM endpoints and RAG configurations.
  • Deploy and manage automation runtime infrastructure: Azure Logic Apps, Power Automate, Azure Functions, Durable Functions, and event-driven triggers for intelligent workflows.
  • Maintain agent hosting environments (Chainlit, FastAPI, Teams bots) for the HR PM Agent and future agentic solutions, with auto-scaling and health monitoring.
  • Create reusable deployment accelerators (Terraform modules, Helm charts, pipeline templates) to reduce time-to-production for each successive initiative.

FinOps, Security & Compliance

  • Drive Azure cost optimization: commitment-tier analysis, right-sizing, automated shutdown policies, and token consumption tracking across LLM endpoints.
  • Implement RBAC, managed identities, Key Vault integration, and least-privilege access across all platform components.
  • Ensure SOX compliance, data residency, and governance using Microsoft Purview, Defender XDR, and Azure Policy.
  • Manage secrets, certificates, API key rotation, and Entra ID integration for platform authentication across global regions.
  • Produce monthly infrastructure cost and performance reports with spend trends, cost-per-query, and optimization metrics.

Requirements

  • 7+ years of cloud infrastructure/DevOps experience with at least 2 years supporting AI/ML, automation, or data platform workloads at scale.
  • Expert-level Azure skills: Databricks, Cosmos DB, Azure Functions, Logic Apps, ADLS Gen2, Azure AI Search, Azure OpenAI Service, Container Apps/AKS, and Azure Monitor.
  • Strong IaC proficiency: Terraform (modules, state, workspaces), Bicep, or ARM templates with environment-templated patterns.
  • Hands-on CI/CD engineering: Azure DevOps, GitHub Actions, container registries, Helm charts, and blue/green/canary deployment automation.
  • Solid Python and Bash skills for infrastructure tooling, automation scripts, and deployment utilities.
  • Deep understanding of Azure networking, security (RBAC, managed identities, Key Vault, Private Endpoints, Azure Policy), and cost management.
  • Experience with containerization (Docker) and orchestration (AKS or Container Apps) for production workload and model serving.
  • Familiarity with AI platform infrastructure: Databricks provisioning, Cosmos DB scaling, AI Search management, and LLM endpoint deployment., * Experience deploying RAG platform infrastructure, vector search clusters, and LLM/SLM serving endpoints in production.
  • Hands-on MLOps: model registries, experiment tracking, automated deployment pipelines, and A/B testing infrastructure.
  • Background in enterprise IT environments with M365, Intune, and Entra ID.
  • Azure certifications: AZ-104, AZ-400, AZ-305.
  • FinOps certification or demonstrated cloud cost optimization experience delivering measurable savings.
  • Experience supporting global operations across AMER, EMEA, and APAC with high-availability requirements., * You are a strong communicator, you are persuasive and clear, blending analytics with experience in decision-making.
  • You do not get flustered easily. You can juggle multiple priorities while balancing urgent requests with shifting timelines and deliverables.
  • You are a team builder. You take the time to understand and develop the strengths of your resources while formulating long-term plans for the growth and success of the team.
  • You are naturally curious and driven toward continual improvement. While you celebrate your successes, you take time to review and analyze campaigns for future learning.

Benefits & conditions

Pulled from the full job description

  • 401(k)
  • Health insurance
  • Paid time off
  • Vision insurance
  • Dental insurance
  • Flexible spending account
  • Life insurance, * Travel: <10%
  • Benefits: Healthcare, Dental Care, Vision Insurance, Life Insurance, Paid Time Off, and Paid Leave Programs
  • Must be eligible to work in the United States
  • Must pass comprehensive background and drug screening, * We offer a competitive compensation package with strong benefits, including medical, dental, and vision insurance, a 401K program, flexible spending accounts - even a cell phone subsidy.
  • We foster a culture of appreciation, including peer-to-peer recognition and rewards programs.
  • Fun is part of our DNA, with events, game nights, happy hours, and barbecues.
  • We’re growing - this is a great time to join and make an impact!

About the company

STACK INFRASTRUCTURE (STACK) provides digital infrastructure to scale the world’s most innovative companies. We are an award-winning industry leader in building, owning, and operating highly efficient, cost-effective wholesale, colocation, and cloud data centers. Each of our national facilities meets or exceeds the highest industry standards in all operational categories of availability, security, connectivity, and physical resilience.

STACK offers the scale and geographic reach that rapidly growing hyperscale and enterprise companies need. The world runs on data. Data runs on STACK.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · WWC 2025

2:56 min

Provisioning a secure container infrastructure with Bicep

Matthias Falkenberg +1 · WWC 2022

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · WWC Europe 2026

Videos

See all

Related articles

See all