AI Platform Engineer

Massachusetts Mutual Life Insurance Company
Boston, MA, United States
2 days ago
Apply on www.careerboard.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Working hours
Regular working hours

Tech stack

Artificial Intelligence Amazon Web Services Audit Trail Microsoft Azure Cloud Computing Cloud Engineering Collaborative Learning Software Design Patterns Programming Tools Identity and Access Management Software Safety Pulumi
+9 more
Graphics Processing Unit (GPU) Delivery Pipeline Large Language Models Model Validation AI Platforms Kubernetes Free and Open-Source Software Machine Learning Operations Terraform

Job description

MassMutual’s AI Platform Engineering team is seeking an impact-driven AI Platform Engineer to serve as the technical anchor of our high-performing team. You will lead the design, deployment, set engineering standards, and drive the most complex platform initiatives from concept through production. The Team

This is a unique opportunity to work on the team that builds and operates the platform powering MassMutual’s AI initiatives. The team operates at the intersection of cloud infrastructure, AI/ML systems, and developer experience-delivering foundational capabilities that shape how the entire organization builds and deploys AI. We partner closely with AI engineering, product, and cloud engineering teams across the enterprise, and we invest in growth through a culture of peer learning, candid feedback, and shared technical standards. This team is defined by a shared commitment to engineering excellence, clear documentation, and the kind of technical leadership that makes hard problems tractable. The Impact

  • Define architectural direction at the AI platform component level-cloud infrastructure, AI serving layers, developer tooling, and reliability strategy-and translate it into a concrete, prioritized roadmap.
  • Lead design on one of the platform’s most critical components-LLM gateway, multi-tenant compute isolation, model serving infrastructure, and enterprise integration patterns. Write ADRs that become the team’s engineering standards.
  • Serve as the first point of technical escalation for hard engineering decisions. Run design reviews, provide deep technical feedback on pull requests, and pair with engineers on the gnarliest problems.
  • Own the technical execution of major platform initiatives end to end-scoping, sequencing work, managing technical risk, and driving to production without losing quality.
  • Lead platform reliability strategy: define SLOs, shape the observability strategy, lead incident reviews, and continuously raise the bar on platform stability and operational maturity.
  • Lead technical design of governance and compliance controls-data residency, access management, audit logging, and AI usage policies-to satisfy enterprise customer requirements and compliance frameworks.
  • Drive cross-team alignment with AI engineering, product, and cloud engineering teams; communicate technical trade-offs clearly and represent platform capabilities to senior stakeholders.
  • Raise the technical craft of the team through thorough design reviews, documentation habits, and hands-on pairing-without managing anyone directly., What to Expect as Part of MassMutual and the Team
  • Regular meetings with the AI Platform Engineering team
  • Focused one-on-one meetings with your manager
  • Networking opportunities including access to Asian, Hispanic/Latinx, African American, women, LGBTQIA+, veteran and disability-focused Business Resource Groups
  • Access to learning content on Degreed and other informational platforms
  • Your ethics and integrity will be valued by a company with a strong and stable ethical business with industry leading pay and benefits

Requirements

  • 5+ years in platform, infrastructure, or SRE, with a track record as a technical lead or staff-level IC with team-wide technical scope.
  • Certified Kubernetes Administrator (CKA), Certified Kubernetes Application Developer (CKAD) or equivalent AWS Certifications.
  • 3+ years experience in cloud-native architecture: Kubernetes at scale, managed cloud services, networking, identity federation, and multi-tenancy patterns across AWS, GCP, or Azure.
  • 3+ years experience of proven ownership of complex, multi-month platform initiatives-driven from whiteboard to production, managing ambiguity and technical risk throughout., * Strong IaC and GitOps fluency: Terraform or Pulumi, ArgoCD, with experience standardizing platform tooling and deployment patterns across engineering teams.
  • Hands-on experience with AI/ML infrastructure: model serving, inference pipelines, GPU resource management, or LLM integration patterns.
  • Experience mentoring engineers and influencing technical direction across teams.
  • Strong written communication: clear design docs, useful ADRs, and the ability to explain architectural decisions to both engineers and non-technical stakeholders.
  • LLM serving at scale-vLLM, Triton, Ray Serve-or experience with AI gateway design patterns.
  • Experience building an internal developer platform (IDP) from scratch, with a product mindset that obsesses over internal developer experience.
  • Strong grasp of AI safety, model evaluation, and governance frameworks.
  • Ability to lead through technical credibility-influencing design decisions, driving alignment, and raising quality without formal authority.
  • FinOps or GPU cost optimization experience across large inference workloads.
  • Open-source contributions to platform or ML infrastructure tooling.
  • Comfort with ambiguity: energized by undefined problem spaces, with a habit of building clarity and shared context where there isn’t any.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerboard.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

2:28 min

Understanding Kubernetes architecture and core cluster components

Marc Nimmerrichter · World Congress 2022

1:34 min

Essential commands for running and testing Terraform configurations

Hennie Francis · LIVE

1:55 min

Contrasting Terraform with Pulumi and cloud-specific tools

Devlin Duldulao · LIVE

1:43 min

Platform engineering as the foundation for scaling AI tools

Julia Kordick Julia Kordick · World Congress 2026 Europe

4:04 min

Overview of Kubernetes operators and custom resource definitions

Philipp Krenn · World Congress 2022

Videos

See all

Related articles

See all