Senior Full-Stack Software Engineer - Global AI Platform

John Hancock
Toronto, United States of America
4 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Languages
English
Experience level
Senior

Job location

Toronto, United States of America

Tech stack

Java
API
Artificial Intelligence
Akka
Amazon Web Services (AWS)
Azure
Cloud Computing
Computer Networks
Continuous Integration
Distributed Systems
Fault Tolerance
Identity and Access Management
Python
Performance Tuning
Software Engineering
Reinforcement Learning
Feature Engineering
Multi-Agent Systems
Backend
AI Platforms
Kubernetes
Machine Learning Operations
Terraform
Stream Processing
Data Pipelines
Docker
Microservices

Job description

Experteer Overview In this role you design and deliver a scalable, secure, cloud-native AI platform that enables enterprise-grade AI and agent-powered solutions. You will work with architects, data scientists, and business leaders to ensure a reliable, high-performance infrastructure that supports continuous learning, experimentation, and governance. You'll balance innovation with regulatory compliance, shaping platforms used for diverse AI workloads and real-time agent interactions. Compensation / Benefits * Build and maintain high-performance, fault-tolerant AI platform services with automation-first delivery * Design and maintain platform infrastructure including hardware, software, and network components * Integrate Akka, AdaptiveML workflows, feature stores, model registries, and A/B experimentation * Implement AI Foundry components for orchestration, feature engineering, deployment, and governance * Develop reusable reference patterns and inner-source components meeting reliability and security standards * Create shared runtimes for multi-agent coordination, state management, and messaging * Design interoperable APIs/SDKs for data scientists and developers * Maintain and improve CI/CD pipelines and developer toolchains for compliant delivery * Evaluate new AI/ML infrastructure capabilities and prototype productivity tools * Develop and operate scalable backend services for high-traffic agent interactions and real-time flows * Use cloud-native tech (containers, orchestration, IaC, CI/CD) for reliable, cost-efficient services * Optimize runtime performance across CPU/GPU/accelerator workloads * Monitor and resolve platform issues, improving bottlenecks and reliability * Ensure compliance and security measures across the platform lifecycle * Collaborate with architects and leaders to build robust platforms across AI capability layers * Develop a holistic understanding of data, tools, and cross-team dependencies * Explore new platform solutions to improve service delivery * Perform peer reviews for code and deliverables to support continuous learning Tasks * 5+ years in software engineering * 3+ years leading AI/ML or distributed systems teams/projects * Strong expertise in Akka and event-driven microservices at scale * Hands-on experience with AI Foundry and AdaptiveML or equivalents * Proficiency in Scala or Java (Akka ecosystem) and Python for ML tooling * Experience with stream processing and data pipelines * Solid MLOps background: model registries, feature stores, ML CI/CD, Docker, Kubernetes * Cloud proficiency (AWS/Azure), Terraform/IaC, and secrets/IAM * Deep understanding of distributed systems: consistency, partitioning, backpressure, resilience * Strong communication and documentation skills * Preferred: online learning, reinforcement learning, or active learning in production * Knowledge of responsible AI, model risk and fairness/bias assessment * Performance optimization for low-latency inference; GPU/accelerator utilization * Experience in regulated industries with audit and governance requirements Key requirements * health, dental, mental health, vision benefits * short- and long-term disability * life and AD&D insurance * adoption/surrogacy and wellness benefits * retirement savings plans with employer matching * paid time off including holidays, vacation, personal, sick days

Requirements

delivery * Perform peer reviews for code and deliverables to support continuous learning Tasks * 5+ years in software engineering * 3+ years leading AI/ML or distributed systems teams/projects * Strong expertise in Akka and event-driven microservices at scale * Hands-on experience with AI Foundry and AdaptiveML or equivalents * Proficiency in Scala or Java (Akka ecosystem) and Python for ML tooling * Experience with stream processing and data pipelines * Solid MLOps background: model registries, feature stores, ML CI/CD, Docker, Kubernetes * Cloud proficiency (AWS/Azure), Terraform/IaC, and secrets/IAM * Deep understanding of distributed systems: consistency, partitioning, backpressure, resilience * Strong communication and documentation skills * Preferred: online learning, reinforcement learning, or active learning in production * Knowledge of responsible AI, model risk and fairness/bias assessment * Performance optimization for low-latency inference; GPU/accelerator aaaaaaaaaz _ * Experience in regulated industries with audit and governance requirements Key requirements * health, dental, mental health, vision benefits * short- and long-term disability * life and AD&D insurance * adoption/surrogacy and wellness benefits * retirement savings plans with employer matching * paid time off including holidays, vacation, personal, sick days

Apply for this position