AI Platform Operations Manager

Stack Infrastructure
Denver, CO, United States
9 days ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$128,260.0 - $146,018.0
Working hours
Regular working hours
Job source

Tech stack

Microsoft Windows Artificial Intelligence Application Layers Application Performance Management Application Release Automation Microsoft Azure Bash Shell Cloud Computing Configuration Management Code Review Cyber Security Continuous Integration
+47 more
Data Centers DevOps Github Monitoring of Systems Python (Programming Language) Key Management Log Analysis Project Management Software Netsuite Open Source Technology Windows PowerShell Release Management Reliability Engineering Prometheus Azure Machine Learning Azure Data Lake Search Technologies Software Deployment Systems Integration Management of Software Versions Policy as Code Data Logging Enterprise Software Applications Cloud Monitoring Autoscaling Large Language Models Grafana Data Strategy Containerization AI Platforms Git Flow Kubernetes Infrastructure Automation Frameworks Information Technology Deployment Automation Bicep Cosmos DB Azure AKS Machine Learning Operations Virtual Agents Terraform Devsecops Workday Docker Key Vault Databricks Vulnerability Analysis

Job description

The DevOps Engineer, AI Platform is responsible for automating, deploying, and operating the infrastructure and delivery pipelines that support STACK’s enterprise AI platform on Azure. This is a hands-on engineering role focused on build and run - not oversight.

Reporting to Head of AI, Enterprise AI & Data Strategy org this individual owns the infrastructure-as-code, CI/CD, containerization, observability, and release automation that allow AI engineers and enterprise application teams to ship agentic AI solutions, RAG pipelines, and integration services reliably and repeatably. The role sits at the intersection of cloud infrastructure, platform engineering, and MLOps - turning platform architecture into automated, governed, observable, and cost-efficient environments that teams across the organization build on., Infrastructure Automation & Infrastructure as Code

  • Build, maintain, and version infrastructure-as-code modules for Azure environments using Terraform, Bicep, or ARM, including compute, networking, storage, identity, and AI platform resources.
  • Automate provisioning of AI platform components - Azure AI Foundry, Azure OpenAI Service, Azure AI Search, Cosmos DB, ADLS Gen2, and Databricks - as reusable, parameterized deployment patterns.
  • Maintain environment parity across development, test, and production, including configuration management, drift detection, and remediation.
  • Implement and enforce tagging, naming, and resource organization standards that support governance, chargeback, and lifecycle management.
  • Automate routine platform operations - patching, certificate rotation, key and secret rotation, backup validation, and disaster recovery testing.

CI/CD & Release Engineering

  • Design, build, and operate CI/CD pipelines in Azure DevOps or GitHub Actions for application code, infrastructure code, container images, and AI/agent deployments.
  • Implement automated build, test, security scanning, artifact management, and promotion gates across environments.
  • Establish branching strategies, code review standards, and release management practices in partnership with AI engineering and enterprise application teams.
  • Build deployment automation for agentic AI services, MCP (Model Context Protocol) servers, and integration workloads running on Azure Container Apps and Azure Kubernetes Service (AKS).
  • Support model and prompt release workflows - versioning, staged rollout, evaluation gates, and rollback procedures for LLM-based applications.

Container Platform & AI Workload Operations

  • Operate and tune AKS and Azure Container Apps, including cluster upgrades, node pool sizing, autoscaling, ingress, networking, and workload isolation.
  • Build and maintain container images, base image standards, and registry governance in Azure Container Registry.
  • Manage compute scheduling and scaling for AI workloads, including GPU-backed and inference-heavy workloads where required.
  • Implement resiliency patterns - health probes, retries, throttling, quota management, and failover - for AI endpoints and integration services.

Observability, Reliability & Incident Response

  • Instrument platform and AI services with logging, metrics, tracing, and alerting using Azure Monitor, Log Analytics, Application Insights, and equivalent open-source tooling.
  • Build dashboards and service-level indicators covering platform availability, latency, throughput, error rates, token consumption, and model endpoint performance.
  • Participate in on-call rotation, lead incident triage and resolution for platform issues, and drive root cause analysis and corrective actions.
  • Develop and maintain runbooks, operational documentation, and automated remediation for recurring issues.

Security, Governance & Cost Optimization

  • Implement DevSecOps practices - secrets management in Azure Key Vault, managed identity usage, least-privilege access, dependency and container vulnerability scanning, and policy-as-code.
  • Partner with Information Security to ensure pipelines and environments meet enterprise security, data residency, and compliance requirements.
  • Support Azure FinOps practices through cost visibility, rightsizing, reserved capacity, and automated controls on non-production and idle resources.
  • Maintain audit trails and change records for infrastructure and release activity.

Delivery & Cross-Functional Collaboration

  • Work directly with AI engineers, data engineers, and enterprise application teams to remove deployment friction and improve time-to-production for AI solutions.
  • Translate platform architecture and standards into automated, self-service capabilities that teams can consume without deep infrastructure knowledge.
  • Contribute to platform engineering standards, reference implementations, and internal documentation.
  • Provide technical escalation support for build, deployment, and environment issues.

Requirements

  • Bachelor’s degree in Computer Science, Information Technology, Engineering, or related field, or equivalent practical experience.
  • 5+ years of hands-on experience in DevOps, site reliability engineering, or platform engineering roles with a strong delivery track record.
  • Strong proficiency in Infrastructure as Code - Terraform, Bicep, or ARM - including module design, state management, and reusable patterns.
  • Proven experience building and operating CI/CD pipelines in Azure DevOps or GitHub Actions.
  • Hands-on experience with containerization and orchestration - Docker, Azure Kubernetes Service (AKS), and Azure Container Apps or equivalent.
  • Solid working knowledge of Azure core services - compute, networking (VNet, NSG, Private Endpoints), storage, identity (Entra ID), and Key Vault.
  • Strong scripting and automation skills in Python, PowerShell, or Bash.
  • Experience with monitoring and observability tooling - Azure Monitor, Log Analytics, Application Insights, Prometheus, or Grafana.
  • Working knowledge of Git-based workflows, code review practices, and artifact/registry management.
  • Demonstrated ability to troubleshoot production issues across infrastructure, network, and application layers., * Microsoft Certified: DevOps Engineer Expert (AZ-400), Azure Administrator (AZ-104), or Certified Kubernetes Administrator (CKA).
  • Experience deploying and operating AI/ML workloads - model endpoints, RAG pipelines, vector databases, or agentic services in production.
  • Familiarity with MLOps tooling and practices - Azure Machine Learning, MLflow, Databricks, or equivalent model lifecycle platforms.
  • Experience deploying MCP (Model Context Protocol) servers or similar integration services connecting AI agents to enterprise systems.
  • Exposure to agentic AI frameworks such as Semantic Kernel, LangGraph, or AutoGen from a deployment and operations perspective.
  • Experience with GPU compute provisioning, quota management, and inference cost optimization.
  • Knowledge of FinOps frameworks and Azure cost optimization practices.
  • Experience integrating with enterprise systems such as Microsoft 365, Freshworks ITSM, Workday, NetSuite, or Procore.
  • Experience in data center, hyperscale, or infrastructure-intensive industry environments., * You are a strong communicator, you are persuasive and clear, blending analytics with experience in decision-making.
  • You do not get flustered easily. You can juggle multiple priorities while balancing urgent requests with shifting timelines and deliverables.
  • You are a team builder. You take the time to understand and develop the strengths of your resources while formulating long-term plans for the growth and success of the team.
  • You are naturally curious and driven toward continual improvement. While you celebrate your successes, you take time to review and analyze campaigns for future learning.

Benefits & conditions

Pulled from the full job description

  • 401(k)
  • Health insurance
  • Paid time off
  • Vision insurance
  • Dental insurance
  • Flexible spending account
  • Life insurance, * Travel: <10%
  • Benefits: Healthcare, Dental Care, Vision Insurance, Life Insurance, Paid Time Off, and Paid Leave Programs
  • Must be eligible to work in the United States
  • Must pass comprehensive background and drug screening, * We offer a competitive compensation package with strong benefits, including medical, dental, and vision insurance, a 401K program, flexible spending accounts - even a cell phone subsidy.
  • We foster a culture of appreciation, including peer-to-peer recognition and rewards programs.
  • Fun is part of our DNA, with events, game nights, happy hours, and barbecues.
  • We’re growing - this is a great time to join and make an impact!

About the company

STACK INFRASTRUCTURE (STACK) provides digital infrastructure to scale the world’s most innovative companies. We are an award-winning industry leader in building, owning, and operating highly efficient, cost-effective wholesale, colocation, and cloud data centers. Each of our national facilities meets or exceeds the highest industry standards in all operational categories of availability, security, connectivity, and physical resilience.

STACK offers the scale and geographic reach that rapidly growing hyperscale and enterprise companies need. The world runs on data. Data runs on STACK.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters · World Congress 2023

2:56 min

Provisioning a secure container infrastructure with Bicep

Matthias Falkenberg +1 · World Congress 2022

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · World Congress 2026 Europe

2:40 min

Using GitHub primitives for internal documentation and corporate operations

Kyle Daigle · Coffee With Developers

Videos

See all

Related articles

See all