Platform Engineer - Senior - US

Quantiphi, Inc.
Boston, MA, United States
15 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Automated Storage and Retrieval Systems Microsoft Azure Computer Clusters Profiling Nvidia CUDA Linux Distributed Computing Environment Interoperability OpenShift Performance Tuning Ansible
+12 more
Google Cloud EHR Systems Fast Healthcare Interoperability Resources Large Language Models Kubernetes Infrastructure Automation Frameworks Health Level Seven International Slurm Machine Learning Operations TensorRT Terraform Oracle Cloud Infrastructure

Job description

  • Design and implement scalable infrastructure for LLM and GenAI workloads across multi-GPU environments
  • Perform GPU profiling, benchmarking, and performance optimization for distributed training workloads
  • Manage and schedule compute-intensive jobs using Slurm-based clusters and OpenShift/Kubernetes environments
  • Enable and optimize the NVIDIA GPU stack (CUDA, cuDNN, NCCL, Triton, RAPIDS, etc.)
  • Collaborate with cross-functional teams to deploy models in research and production environments
  • Build and support GenAI pipelines (fine-tuning, RAG, multi-modal inferencing, LLMOps)
  • Develop reusable infrastructure templates using tools like Terraform and Helm
  • Contribute to internal innovation (PoCs, workshops) and support client-facing delivery engagements

Requirements

We are looking for a highly skilled Architect - Platform Engineer to design, optimize, and scale infrastructure for GenAI and LLM workloads. This role is ideal for someone with deep hands-on experience in GPU profiling, distributed training, and high-performance compute environments., * Strong experience with Slurm and distributed training environments

  • Hands-on expertise with Red Hat OpenShift and/or Kubernetes
  • Deep knowledge of the NVIDIA GPU ecosystem (CUDA, cuDNN, NCCL, Nsight, Triton/TensorRT)
  • Strong foundation in Linux systems, performance tuning, and multi-GPU optimization
  • Experience deploying GenAI workloads (LLM fine-tuning, RAG pipelines, multi-modal systems)
  • Familiarity with Infrastructure-as-Code tools (Terraform, Ansible)
  • Experience with cloud GPU environments (Google Cloud Platform, Azure, AWS, OCI) and/or on-prem GPU clusters

Other Qualifications (OQs):

  • Experience with NVIDIA NIMs, DGX systems, or GPU-accelerated containers
  • Knowledge of LLMOps frameworks and MLOps integration
  • Familiarity with vector databases and retrieval systems for RAG architectures
  • Comfortable working in client-facing environments and collaborating with AI solution teams

Healthcare Domain Experience (Nice to Have):

  • Experience working with FHIR R4, HL7 v2, or SMART on FHIR
  • Integration with EHR systems (e.g., Epic)
  • Understanding of HIPAA compliance and healthcare data privacy
  • Exposure to clinical workflows, CDS Hooks, or patient-facing applications
  • Experience building clinical decision support systems or healthcare interoperability solutions

About the company

While technology is the heart of our business, a global and diverse culture is the heart of our success. We love our people and we take pride in catering them to a culture built on transparency, diversity, integrity, learning and growth.

If working in an environment that encourages you to innovate and excel, not just in professional but personal life, interests you- you would enjoy your career with Quantiphi!

About Quantiphi:

Quantiphi is an award-winning, AI-First global digital engineering company that helps the world’s leading Fortune 1000 organizations transform bold ideas into measurable business impact. We go beyond building innovative AI technologies, we solve the problems that matter most to our clients.

Since our founding in 2013, Quantiphi has built a proven track record of turning complex challenges into meaningful outcomes across industries.

Headquartered in Boston, with more than 4,000 professionals worldwide, we partner with global enterprises to deliver large-scale digital, cloud, and AI-driven transformation. #SolvingWhatMatters

We are an Elite and Premier partner to Google Cloud, AWS, NVIDIA, Snowflake, and other leading technology platforms, and our work has been recognized across the industry, including:

  • 21 Google Cloud Partner of the Year awards in the past 10 years
  • 3 AWS AI/ML Partner of the Year awards
  • 3 NVIDIA Partner of the Year awards
  • 3 Snowflake Partner of the Year awards
  • Rated Leaders by Gartner, Forrester, IDC, ISG, Everest Group and other leading analyst firms

Quantiphi delivers First-in-class AI solutions across Life Sciences, Healthcare, Banking, Financial Services, CPG, Manufacturing, Energy, High-Tech, Telecommunications, etc., powered by cutting-edge Generative AI and Agentic AI accelerators.

We are also proud to be certified as a Great Place to Work, reflecting our commitment to our people and our culture., What’s in it for YOU at Quantiphi:

  • Make an impact at one of the world’s fastest-growing AI-first digital engineering companies.
  • Upskill and discover your potential as you solve complex challenges in cutting-edge areas of technology alongside passionate, talented colleagues.
  • Work where innovation happens - work with disruptive innovators in a research-focused organization with 60+ patents filed across various disciplines.
  • Stay ahead of the curve, immerse yourself in breakthrough AI, ML, data, and cloud technologies and gain exposure working with Fortune 500 companies.

If you like wild growth and working with happy, enthusiastic over-achievers, you’ll enjoy your career with us!

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:42 min

Automating Skupper deployments using Ansible

Alex Soto Alex Soto · World Congress 2024

2:22 min

Infrastructure barriers and compliance risks in research

Jeremy Murray Jeremy Murray · World Congress 2026 Europe

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

2:08 min

Essential engineering roles in the generative AI space

Mary Grygleski Mary Grygleski · LIVE

3:19 min

Executing complex workflows using Ansible Automation Platform

Goetz Rieger Goetz Rieger · World Congress 2025

1:51 min

Managing GPU quotas and multi-tenancy with Kueue

Jeremy Murray Jeremy Murray · World Congress 2026 Europe

Videos

See all

Related articles

See all