> Markdown version of [/jobs/ext/1959260-lead-platform-reliability-engineer-global-ai-platform-solutions](https://www.wearedevelopers.com/jobs/ext/1959260-lead-platform-reliability-engineer-global-ai-platform-solutions). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Lead Platform Reliability Engineer, Global AI Platform & Solutions - **Company:** John Hancock - **Location:** Toronto, OH, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Artificial Intelligence, Audit Trail, Microsoft Azure, Cloud Computing, Continuous Integration, Data Systems, DevOps, Distributed Systems, Python (Programming Language), Key Management, Performance Tuning, Role-Based Access Control, Reliability Engineering, Ansible, TypeScript, Data Logging, System Availability, Large Language Models, Concurrency, Mttr, Backend, AI Platforms, Kubernetes, Information Technology, Machine Learning Operations, Api Design, Terraform, Automation Anywhere - **Published:** August 6, 2026 - **Apply:** https://us.experteer.com/career/view-jobs/lead-platform-reliability-engineer-global-ai-platform-and-solutions-toronto-oh-usa-58815870 ## About the Role _ prevent configuration drift * Enforce identity/RBAC, secrets management, supply chain security, regulatory controls; collaborate with risk and audit * Optimize resource usage and cost: capacity planning, rightsizing, reservations/spot * Safe change management: progressive delivery, policy-as-code guardrails * Treat platform as a product: define operational SLAs aligned to roadmap, service catalog, and developer experience * Collaborate with global engineering, security, and AI governance teams for cross-geo regulations and data residency * Operate scalable backend services for high-traffic agent interactions, retrieval operations, and real-time execution * Maintain AI services runbooks and playbooks; enable GOCC Tasks * Bachelor's in Computer Science/Engineering or equivalent (not strictly required) * 5-8 years in DevOps/Platform Engineering or Production Operations * Proven track record with large-scale distributed systems and on-call experience * Cloud-native development experience: Azure, Kubernetes, containers, CI/CD, observability stacks * Proficiency in Python and/or Java/Scala/TypeScript for backend services and automation * Knowledge of AI solutions, LLM systems, retrieval architectures, embeddings, vector stores, prompt/tool orchestration, and agent workflows * Understanding of API design, asynchronous workflows, concurrency, reliability engineering (SLOs, error budgets), and performance tuning * Familiarity with security, governance, and compliance for AI/data systems (authN/authZ, data protection, audit logging, model governance) * Ability to collaborate across global teams and translate business requirements into platform capabilities and operational SLAs Key requirements * health and dental insurance * mental health support * vision coverage * disability and life insurance * retirement savings plans * paid time off ## Description Experteer Overview In this role you will lead reliability for a shared platform that powers internal AI solution development. You'll partner with global teams to define and meet SLOs, scale platforms, and improve observability and incident response. You'll build self-service and automation to reduce toil, while ensuring security, governance, and data residency compliance. This is a hands-on, cross-functional role shaping a platform-as-a-product that enables high-traffic AI workflows at scale. Compensation / Benefits * Define SLOs/SLIs, manage operations budgets, reduce MTTR, capacity planning, and autoscale tuning * Build and maintain logging, metrics, tracing, alerting; instrument platform components; create runbooks and dashboards * On-call incident response: triage, mitigate, root-cause, postmortems and corrective actions * Develop self-service capabilities and automation: AIOps/MLOps/GitOps/CI/CD pipelines and operational automations * Manage infrastructure via Terraform/Ansible; prevent configuration drift * Enforce identity/RBAC, secrets management, supply chain security, regulatory controls; collaborate with risk and audit * Optimize resource usage and cost: capacity planning, rightsizing, reservations/spot * Safe change management: progressive delivery, policy-as-code guardrails * Treat platform as a product: define operational SLAs aligned to roadmap, service catalog, and developer experience * Collaborate with global engineering, security, and AI governance teams for cross-geo regulations and data residency * Operate scalable backend services for high-traffic agent interactions, retrieval operations, and real-time execution * Maintain AI services runbooks and playbooks; enable GOCC Tasks * Bachelor's in Computer Science/Engineering or equivalent (not strictly required) * 5-8 years in DevOps/Platform Engineering or Production Operations * Proven track record with large-scale distributed systems and on-call experience * Cloud-native development experience: Azure, Kubernetes, containers, CI/CD, observability stacks * Proficiency in Python and/or Java/Scala/TypeScript for backend services and automation * Knowledge of AI solutions, LLM systems, retrieval architectures, embeddings, vector stores, prompt/tool orchestration, and agent workflows * Understanding of API design, asynchronous workflows, concurrency, reliability engineering (SLOs, error budgets), and performance tuning * Familiarity with security, governance, and compliance for AI/data systems (authN/authZ, data protection, audit logging, model governance) * Ability to collaborate across global teams and translate business requirements into platform capabilities and operational SLAs Key requirements * health and dental insurance * mental health support * vision coverage * disability and life insurance * retirement savings plans * paid time off ## Related Videos - [What Developers Get Wrong About Application Quality](https://www.wearedevelopers.com/videos/233-what-developers-get-wrong-about-application-quality) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Dev & Test in the Cloud? Deploy your cloud environments with Ansible & Terraform](https://www.wearedevelopers.com/videos/1607-dev-test-in-the-cloud-deploy-your-cloud-environments-with-ansible-terraform) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [#90DaysOfDevOps - The DevOps Learning Journey](https://www.wearedevelopers.com/videos/548-90daysofdevops-the-devops-learning-journey) - [AI-Augmented DevOps with Platform Engineering](https://www.wearedevelopers.com/videos/1614-ai-augmented-devops-with-platform-engineering) ## Related Articles - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this)