> Markdown version of [/jobs/ext/3095214-principal-cloud-engineer-ai](https://www.wearedevelopers.com/jobs/ext/3095214-principal-cloud-engineer-ai). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal Cloud Engineer AI - **Company:** BMO Financial Group - **Location:** San Francisco, CA, United States - **Experience:** Expert - **Salary:** $120,000.0 - $250,000.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Amazon Web Services, Microsoft Azure, Cloud Computing, Cloud Engineering, Computer Clusters, Nvidia CUDA, Databases, Continuous Integration, Identity and Access Management, Python (Programming Language), Key Management, Network Security, Machine Learning, Peering, Microsoft Platform Builder, Prometheus, Azure Machine Learning, Azure Data Lake, Data Streaming, TypeScript, AI Infrastructure, Policy as Code, Data Logging, Real Time Systems, Autoscaling, Istio, Large Language Models, Grafana, Multi-Cloud, Caching, Backend, Event Driven Architecture, Data Lakes, AI Platforms, Kubernetes, HuggingFace, Bicep, Apache Kafka, Machine Learning Operations, Front End Software Development, TensorRT, Terraform, Devsecops, Serverless Computing, Key Vault, Microservices - **Published:** September 26, 2026 - **Apply:** https://dejobs.org/x/x/202231CBFA11424BA47332E9A3FB0FB8/job/ ## About the Role * Bachelor's/Master's/PhD in CS, Engineering, or related field * 7+ years building large-scale distributed cloud infrastructure * 5+ years hands-on with Azure/AWS * Proven experience with AI/ML infra: GPU clusters, Kubernetes, CI/CD, observability * Strong in IaC (Terraform/Bicep), Kubernetes, networking, security * Expertise in cloud-native patterns: containers, service mesh, serverless * Familiarity with MLOps/LLMOps infra: model serving, feature stores, vector DBs * Programming in Python (infra automation) and one of Go/TypeScript for tooling * Understanding of frontend/backend integration for AI services * Familiarity with MLOps/LLMOps infra: model serving, feature stores, vector DBs * Programming in Python (infra automation) and one of Go/TypeScript for tooling * Understanding of frontend/backend integration for AI services Nice-to-Have * GPU optimization (CUDA/NCCL, TensorRT-LLM) * Observability tools (Prometheus, Grafana, OpenTelemetry) * Event streaming (Kafka/Azure Event Hubs), real-time systems * Experience with AI platform products (Azure ML, MLflow, KServe, Hugging Face) Success Metrics * Reliability & Performance: SLOs met for infra services, GPU utilization optimized * Security & Compliance: Zero critical findings, auditable infra * Cost Efficiency: Reduced GPU/infra spend via FinOps strategies * Developer Velocity: Faster provisioning and deployment of AI infra * Technical Leadership: Influence on infra standards, mentorship, reusable patterns ## Description The Team - We accelerate BMO's AI journey by building cloud-native AI solutions. Our team combines engineering excellence with cutting-edge AI to deliver scalable, secure, and responsible solutions that power business innovation across the bank. We enable and accelerate our partners on their AI journeys across the enterprise, helping teams across BMO unlock value at scale. We support one another in times of need and take pride in our work. We are engineers, AI practitioners, platform builders, thought leaders, multipliers, and coders. Above all, we are a global team of diverse individuals who enjoy working together to create smart, secure, and scalable solutions that make an impact across the enterprise. Our ambition is bold: deploy our capital and resources to their highest and most profitable use through a digital-first operating model, powered by data and AI-driven decisions. The Impact As a Principal AI & Cloud Engineer , you are a hands-on technical developer who designs, builds, and scales cloud-native AI solutions and products. You help set engineering standards, establish patterns, mentor senior engineers, and partner with multiple teams to deliver resilient, governed, and cost-efficient AI at enterprise scale. You'll help shape and evolve our AI cloud strategy from model serving and LLMOps to security, observability, and compliance so teams across the bank can innovate safely and rapidly. You will advance BMO's Digital First strategy by: * Defining reference and production-grade solutions for AI/GenAI on cloud (Azure/AWS preferred; multi-cloud aware). * Building reusable, secure, and observable components (APIs, SDKs, microservices, pipelines). * Operationalizing LLMs and RAG with strong controls and Responsible AI guardrails. * Driving platform roadmaps that enable faster delivery, lower risk, and measurable business outcomes. What's In It for You * Influence the technical direction of AI and the platform primitives others build on. * Ship high-impact systems used across many business lines and products. * Work across the full stack: cloud infra, data/feature pipelines, model serving, LLMOps, and DevSecOps. * Partner with a leadership team invested in your growth and thought leadership. Responsibilities Infrastructure & Platform Builder * Design, build, and operate cloud-native AI infrastructure for ML/GenAI workloads: * Compute: GPU/CPU clusters, autoscaling, spot instance strategies * Networking: Azure VNet, Private Link, peering, multi-region HA/DR * Storage & Databases: high-performance data lakes (e.g., Azure Data Lake Storage) , relational DBs, vector DBs (FAISS, Milvus, Pinecone, pgvector) * Security: IAM, Key Vault-backed secrets management, encryption, policy-as-code * Implement observability and reliability for AI infra: * Metrics (latency, throughput, GPU utilization, cost) * Logging/tracing (OpenTelemetry), SLOs/SLIs for infra services * Build CI/CD and GitOps pipelines for infrastructure-as-code (Terraform/Bicep) and AI platform components * Drive FinOps for AI infra: GPU rightsizing, caching, inference optimization, cost governance Application & Service Enablement * Enable frontend and backend services for AI platforms: * Secure APIs, microservices, and event-driven architectures * Integration with custom model runtimes (TensorRT-LLM, vLLM, Triton/KServe) * Provide infrastructure support for RAG systems: embeddings, chunking, retrieval pipelines * Ensure scalable serving infrastructure for LLMs and ML models with caching and token optimization Strategy & Architecture * Define and evolve AI infrastructure reference architecture for cloud (Azure preferred): * Container orchestration (Kubernetes), service mesh, ingress * Serverless/event-driven patterns for AI pipelines * Multi-region, HA/DR, compliance-ready designs * Establish standards and best practices for containerization, IaC, and secure networking for AI systems Security, Risk & Governance * Implement defense-in-depth for AI infra: * IAM least privilege, private networking, KMS/Key Vault, SBOM, image signing * Ensure compliance and Responsible AI controls at infra level: * Data residency, encryption, lineage, audit readiness Delivery & Operations * Lead infrastructure discovery and solution design with stakeholders * Operate platforms with SRE principles: error budgets, incident response, chaos testing * Mentor engineers; create reusable IaC modules, templates, and golden paths ## Related Videos - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Back(end) to the Future: Embracing the continuous Evolution of Infrastructure and Code](https://www.wearedevelopers.com/videos/440-back-end-to-the-future-embracing-the-continuous-evolution-of-infrastructure-and-code) - [Rate-limiting using eBPF and Istio: How to protect your SaaS customers from themselves](https://www.wearedevelopers.com/videos/100220-rate-limiting-using-ebpf-and-istio-how-to-protect-your-saas-customers-from-themselves) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) - [Introduction to Azure Machine Learning](https://www.wearedevelopers.com/videos/368-introduction-to-azure-machine-learning) ## Related Articles - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it)