> Markdown version of [/jobs/ext/2569205-platform-engineer-senior-us](https://www.wearedevelopers.com/jobs/ext/2569205-platform-engineer-senior-us). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Platform Engineer - Senior - US - **Company:** Quantiphi, Inc. - **Location:** Boston, MA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Automated Storage and Retrieval Systems, Microsoft Azure, Computer Clusters, Profiling, Nvidia CUDA, Linux, Distributed Computing Environment, Interoperability, OpenShift, Performance Tuning, Ansible, Google Cloud, EHR Systems, Fast Healthcare Interoperability Resources, Large Language Models, Kubernetes, Infrastructure Automation Frameworks, Health Level Seven International, Slurm, Machine Learning Operations, TensorRT, Terraform, Oracle Cloud Infrastructure - **Published:** August 20, 2026 - **Apply:** https://www.dice.com/job-detail/e4c6876d-2952-48cc-98f8-9a19c950c22a ## About the Role We are looking for a highly skilled Architect - Platform Engineer to design, optimize, and scale infrastructure for GenAI and LLM workloads. This role is ideal for someone with deep hands-on experience in GPU profiling, distributed training, and high-performance compute environments., * Strong experience with Slurm and distributed training environments * Hands-on expertise with Red Hat OpenShift and/or Kubernetes * Deep knowledge of the NVIDIA GPU ecosystem (CUDA, cuDNN, NCCL, Nsight, Triton/TensorRT) * Strong foundation in Linux systems, performance tuning, and multi-GPU optimization * Experience deploying GenAI workloads (LLM fine-tuning, RAG pipelines, multi-modal systems) * Familiarity with Infrastructure-as-Code tools (Terraform, Ansible) * Experience with cloud GPU environments (Google Cloud Platform, Azure, AWS, OCI) and/or on-prem GPU clusters Other Qualifications (OQs): * Experience with NVIDIA NIMs, DGX systems, or GPU-accelerated containers * Knowledge of LLMOps frameworks and MLOps integration * Familiarity with vector databases and retrieval systems for RAG architectures * Comfortable working in client-facing environments and collaborating with AI solution teams Healthcare Domain Experience (Nice to Have): * Experience working with FHIR R4, HL7 v2, or SMART on FHIR * Integration with EHR systems (e.g., Epic) * Understanding of HIPAA compliance and healthcare data privacy * Exposure to clinical workflows, CDS Hooks, or patient-facing applications * Experience building clinical decision support systems or healthcare interoperability solutions ## Description * Design and implement scalable infrastructure for LLM and GenAI workloads across multi-GPU environments * Perform GPU profiling, benchmarking, and performance optimization for distributed training workloads * Manage and schedule compute-intensive jobs using Slurm-based clusters and OpenShift/Kubernetes environments * Enable and optimize the NVIDIA GPU stack (CUDA, cuDNN, NCCL, Triton, RAPIDS, etc.) * Collaborate with cross-functional teams to deploy models in research and production environments * Build and support GenAI pipelines (fine-tuning, RAG, multi-modal inferencing, LLMOps) * Develop reusable infrastructure templates using tools like Terraform and Helm * Contribute to internal innovation (PoCs, workshops) and support client-facing delivery engagements ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Dev & Test in the Cloud? Deploy your cloud environments with Ansible & Terraform](https://www.wearedevelopers.com/videos/1607-dev-test-in-the-cloud-deploy-your-cloud-environments-with-ansible-terraform) - [LLMOps-driven fine-tuning, evaluation, and inference with NVIDIA NIM & NeMo Microservices](https://www.wearedevelopers.com/videos/1582-llmops-driven-fine-tuning-evaluation-and-inference-with-nvidia-nim-nemo-microservices) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) - [Your Next AI Needs 10,000 GPUs. Now What?](https://www.wearedevelopers.com/videos/1590-your-next-ai-needs-10-000-gpus-now-what) ## Related Articles - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [The Fastest-Growing Tech Sectors to Look Out for in 2025](https://www.wearedevelopers.com/magazine/373-the-fastest-growing-tech-sectors-to-look-out-for-in-2025) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development)