Senior Full-Stack Software Engineer - Global AI Platform

John Hancock
Toronto, OH, United States
5 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
3 years minimum
Working hours
Regular working hours

Tech stack

Java (Programming Language) Application Programming Interfaces (APIs) Artificial Intelligence Akka (Toolkit) Amazon Web Services Microsoft Azure Cloud Computing Computer Networks Continuous Integration Distributed Systems Fault Tolerance Identity and Access Management
+15 more
Python (Programming Language) Performance Tuning Software Engineering Reinforcement Learning Feature Engineering Multi-Agent Systems Backend AI Platforms Kubernetes Machine Learning Operations Terraform Stream Processing Data Pipelines Docker Microservices

Job description

Experteer Overview In this role you design and deliver a scalable, secure, cloud-native AI platform that enables enterprise-grade AI and agent-powered solutions. You will work with architects, data scientists, and business leaders to ensure a reliable, high-performance infrastructure that supports continuous learning, experimentation, and governance. You’ll balance innovation with regulatory compliance, shaping platforms used for diverse AI workloads and real-time agent interactions. Compensation / Benefits * Build and maintain high-performance, fault-tolerant AI platform services with automation-first delivery * Design and maintain platform infrastructure including hardware, software, and network components * Integrate Akka, AdaptiveML workflows, feature stores, model registries, and A/B experimentation * Implement AI Foundry components for orchestration, feature engineering, deployment, and governance * Develop reusable reference patterns and inner-source components meeting reliability and security standards * Create shared runtimes for multi-agent coordination, state management, and messaging * Design interoperable APIs/SDKs for data scientists and developers * Maintain and improve CI/CD pipelines and developer toolchains for compliant delivery * Evaluate new AI/ML infrastructure capabilities and prototype productivity tools * Develop and operate scalable backend services for high-traffic agent interactions and real-time flows * Use cloud-native tech (containers, orchestration, IaC, CI/CD) for reliable, cost-efficient services * Optimize runtime performance across CPU/GPU/accelerator workloads * Monitor and resolve platform issues, improving bottlenecks and reliability * Ensure compliance and security measures across the platform lifecycle * Collaborate with architects and leaders to build robust platforms across AI capability layers * Develop a holistic understanding of data, tools, and cross-team dependencies * Explore new platform solutions to improve service delivery * Perform peer reviews for code and deliverables to support continuous learning Tasks * 5+ years in software engineering * 3+ years leading AI/ML or distributed systems teams/projects * Strong expertise in Akka and event-driven microservices at scale * Hands-on experience with AI Foundry and AdaptiveML or equivalents * Proficiency in Scala or Java (Akka ecosystem) and Python for ML tooling * Experience with stream processing and data pipelines * Solid MLOps background: model registries, feature stores, ML CI/CD, Docker, Kubernetes * Cloud proficiency (AWS/Azure), Terraform/IaC, and secrets/IAM * Deep understanding of distributed systems: consistency, partitioning, backpressure, resilience * Strong communication and documentation skills * Preferred: online learning, reinforcement learning, or active learning in production * Knowledge of responsible AI, model risk and fairness/bias assessment * Performance optimization for low-latency inference; GPU/accelerator utilization * Experience in regulated industries with audit and governance requirements Key requirements * health, dental, mental health, vision benefits * short- and long-term disability * life and AD&D insurance * adoption/surrogacy and wellness benefits * retirement savings plans with employer matching * paid time off including holidays, vacation, personal, sick days

Requirements

delivery * Perform peer reviews for code and deliverables to support continuous learning Tasks * 5+ years in software engineering * 3+ years leading AI/ML or distributed systems teams/projects * Strong expertise in Akka and event-driven microservices at scale * Hands-on experience with AI Foundry and AdaptiveML or equivalents * Proficiency in Scala or Java (Akka ecosystem) and Python for ML tooling * Experience with stream processing and data pipelines * Solid MLOps background: model registries, feature stores, ML CI/CD, Docker, Kubernetes * Cloud proficiency (AWS/Azure), Terraform/IaC, and secrets/IAM * Deep understanding of distributed systems: consistency, partitioning, backpressure, resilience * Strong communication and documentation skills * Preferred: online learning, reinforcement learning, or active learning in production * Knowledge of responsible AI, model risk and fairness/bias assessment * Performance optimization for low-latency inference; GPU/accelerator aaaaaaaaaz _ * Experience in regulated industries with audit and governance requirements Key requirements * health, dental, mental health, vision benefits * short- and long-term disability * life and AD&D insurance * adoption/surrogacy and wellness benefits * retirement savings plans with employer matching * paid time off including holidays, vacation, personal, sick days

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on us.experteer.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:14 min

Exploring internal AI product initiatives and global engineering roles

Maria Apazoglou · Coffee With Developers

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · WWC 2025

1:52 min

Structuring and scaling the backend engineering team

Stefan Lingler Stefan Lingler +1 · Coffee With Developers

1:34 min

Essential commands for running and testing Terraform configurations

Hennie Francis · LIVE

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · WWC Europe 2026

Videos

See all

Related articles

See all