Lead Platform Reliability Engineer, Global AI Platform & Solutions

John Hancock
Toronto, United States of America
4 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Languages
English
Experience level
Senior

Job location

Toronto, United States of America

Tech stack

Java
Artificial Intelligence
Audit Trail
Azure
Cloud Computing
Continuous Integration
Data Systems
DevOps
Distributed Systems
Python
Key Management
Performance Tuning
Role-Based Access Control
Reliability Engineering
Ansible
TypeScript
Data Logging
System Availability
Large Language Models
Concurrency
Mttr
Backend
AI Platforms
Kubernetes
Information Technology
Machine Learning Operations
Api Design
Terraform
Automation Anywhere

Job description

Experteer Overview In this role you will lead reliability for a shared platform that powers internal AI solution development. You'll partner with global teams to define and meet SLOs, scale platforms, and improve observability and incident response. You'll build self-service and automation to reduce toil, while ensuring security, governance, and data residency compliance. This is a hands-on, cross-functional role shaping a platform-as-a-product that enables high-traffic AI workflows at scale. Compensation / Benefits * Define SLOs/SLIs, manage operations budgets, reduce MTTR, capacity planning, and autoscale tuning * Build and maintain logging, metrics, tracing, alerting; instrument platform components; create runbooks and dashboards * On-call incident response: triage, mitigate, root-cause, postmortems and corrective actions * Develop self-service capabilities and automation: AIOps/MLOps/GitOps/CI/CD pipelines and operational automations * Manage infrastructure via Terraform/Ansible; prevent configuration drift * Enforce identity/RBAC, secrets management, supply chain security, regulatory controls; collaborate with risk and audit * Optimize resource usage and cost: capacity planning, rightsizing, reservations/spot * Safe change management: progressive delivery, policy-as-code guardrails * Treat platform as a product: define operational SLAs aligned to roadmap, service catalog, and developer experience * Collaborate with global engineering, security, and AI governance teams for cross-geo regulations and data residency * Operate scalable backend services for high-traffic agent interactions, retrieval operations, and real-time execution * Maintain AI services runbooks and playbooks; enable GOCC Tasks * Bachelor's in Computer Science/Engineering or equivalent (not strictly required) * 5-8 years in DevOps/Platform Engineering or Production Operations * Proven track record with large-scale distributed systems and on-call experience * Cloud-native development experience: Azure, Kubernetes, containers, CI/CD, observability stacks * Proficiency in Python and/or Java/Scala/TypeScript for backend services and automation * Knowledge of AI solutions, LLM systems, retrieval architectures, embeddings, vector stores, prompt/tool orchestration, and agent workflows * Understanding of API design, asynchronous workflows, concurrency, reliability engineering (SLOs, error budgets), and performance tuning * Familiarity with security, governance, and compliance for AI/data systems (authN/authZ, data protection, audit logging, model governance) * Ability to collaborate across global teams and translate business requirements into platform capabilities and operational SLAs Key requirements * health and dental insurance * mental health support * vision coverage * disability and life insurance * retirement savings plans * paid time off

Requirements

_ prevent configuration drift * Enforce identity/RBAC, secrets management, supply chain security, regulatory controls; collaborate with risk and audit * Optimize resource usage and cost: capacity planning, rightsizing, reservations/spot * Safe change management: progressive delivery, policy-as-code guardrails * Treat platform as a product: define operational SLAs aligned to roadmap, service catalog, and developer experience * Collaborate with global engineering, security, and AI governance teams for cross-geo regulations and data residency * Operate scalable backend services for high-traffic agent interactions, retrieval operations, and real-time execution * Maintain AI services runbooks and playbooks; enable GOCC Tasks * Bachelor's in Computer Science/Engineering or equivalent (not strictly required) * 5-8 years in DevOps/Platform Engineering or Production Operations * Proven track record with large-scale distributed systems and on-call experience * Cloud-native development experience: Azure, Kubernetes, containers, CI/CD, observability stacks * Proficiency in Python and/or Java/Scala/TypeScript for backend services and automation * Knowledge of AI solutions, LLM systems, retrieval architectures, embeddings, vector stores, prompt/tool orchestration, and agent workflows * Understanding of API design, asynchronous workflows, concurrency, reliability engineering (SLOs, error budgets), and performance tuning * Familiarity with security, governance, and compliance for AI/data systems (authN/authZ, data protection, audit logging, model governance) * Ability to collaborate across global teams and translate business requirements into platform capabilities and operational SLAs Key requirements * health and dental insurance * mental health support * vision coverage * disability and life insurance * retirement savings plans * paid time off

Apply for this position