Lead Platform Reliability Engineer, Global AI Platform & Solutions

John Hancock
Toronto, OH, United States
5 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours

Tech stack

Java (Programming Language) Artificial Intelligence Audit Trail Microsoft Azure Cloud Computing Continuous Integration Data Systems DevOps Distributed Systems Python (Programming Language) Key Management Performance Tuning
+17 more
Role-Based Access Control Reliability Engineering Ansible TypeScript Data Logging System Availability Large Language Models Concurrency Mttr Backend AI Platforms Kubernetes Information Technology Machine Learning Operations Api Design Terraform Automation Anywhere

Job description

Experteer Overview In this role you will lead reliability for a shared platform that powers internal AI solution development. You’ll partner with global teams to define and meet SLOs, scale platforms, and improve observability and incident response. You’ll build self-service and automation to reduce toil, while ensuring security, governance, and data residency compliance. This is a hands-on, cross-functional role shaping a platform-as-a-product that enables high-traffic AI workflows at scale. Compensation / Benefits * Define SLOs/SLIs, manage operations budgets, reduce MTTR, capacity planning, and autoscale tuning * Build and maintain logging, metrics, tracing, alerting; instrument platform components; create runbooks and dashboards * On-call incident response: triage, mitigate, root-cause, postmortems and corrective actions * Develop self-service capabilities and automation: AIOps/MLOps/GitOps/CI/CD pipelines and operational automations * Manage infrastructure via Terraform/Ansible; prevent configuration drift * Enforce identity/RBAC, secrets management, supply chain security, regulatory controls; collaborate with risk and audit * Optimize resource usage and cost: capacity planning, rightsizing, reservations/spot * Safe change management: progressive delivery, policy-as-code guardrails * Treat platform as a product: define operational SLAs aligned to roadmap, service catalog, and developer experience * Collaborate with global engineering, security, and AI governance teams for cross-geo regulations and data residency * Operate scalable backend services for high-traffic agent interactions, retrieval operations, and real-time execution * Maintain AI services runbooks and playbooks; enable GOCC Tasks * Bachelor’s in Computer Science/Engineering or equivalent (not strictly required) * 5-8 years in DevOps/Platform Engineering or Production Operations * Proven track record with large-scale distributed systems and on-call experience * Cloud-native development experience: Azure, Kubernetes, containers, CI/CD, observability stacks * Proficiency in Python and/or Java/Scala/TypeScript for backend services and automation * Knowledge of AI solutions, LLM systems, retrieval architectures, embeddings, vector stores, prompt/tool orchestration, and agent workflows * Understanding of API design, asynchronous workflows, concurrency, reliability engineering (SLOs, error budgets), and performance tuning * Familiarity with security, governance, and compliance for AI/data systems (authN/authZ, data protection, audit logging, model governance) * Ability to collaborate across global teams and translate business requirements into platform capabilities and operational SLAs Key requirements * health and dental insurance * mental health support * vision coverage * disability and life insurance * retirement savings plans * paid time off

Requirements

_ prevent configuration drift * Enforce identity/RBAC, secrets management, supply chain security, regulatory controls; collaborate with risk and audit * Optimize resource usage and cost: capacity planning, rightsizing, reservations/spot * Safe change management: progressive delivery, policy-as-code guardrails * Treat platform as a product: define operational SLAs aligned to roadmap, service catalog, and developer experience * Collaborate with global engineering, security, and AI governance teams for cross-geo regulations and data residency * Operate scalable backend services for high-traffic agent interactions, retrieval operations, and real-time execution * Maintain AI services runbooks and playbooks; enable GOCC Tasks * Bachelor’s in Computer Science/Engineering or equivalent (not strictly required) * 5-8 years in DevOps/Platform Engineering or Production Operations * Proven track record with large-scale distributed systems and on-call experience * Cloud-native development experience: Azure, Kubernetes, containers, CI/CD, observability stacks * Proficiency in Python and/or Java/Scala/TypeScript for backend services and automation * Knowledge of AI solutions, LLM systems, retrieval architectures, embeddings, vector stores, prompt/tool orchestration, and agent workflows * Understanding of API design, asynchronous workflows, concurrency, reliability engineering (SLOs, error budgets), and performance tuning * Familiarity with security, governance, and compliance for AI/data systems (authN/authZ, data protection, audit logging, model governance) * Ability to collaborate across global teams and translate business requirements into platform capabilities and operational SLAs Key requirements * health and dental insurance * mental health support * vision coverage * disability and life insurance * retirement savings plans * paid time off

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on us.experteer.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

1:42 min

Automating Skupper deployments using Ansible

Alex Soto Alex Soto · WWC 2024

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

3:08 min

Aligning engineering processes with core business impact metrics

Chris Riley · WWC 2021

3:19 min

Executing complex workflows using Ansible Automation Platform

Goetz Rieger Goetz Rieger · WWC 2025

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

Videos

See all

Related articles

See all