> Markdown version of [/jobs/ext/1919574-principal-solution-architect-ai-infrastructure-and-cloud](https://www.wearedevelopers.com/jobs/ext/1919574-principal-solution-architect-ai-infrastructure-and-cloud). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal Solution Architect - AI Infrastructure and Cloud - **Company:** IREN LLC - **Location:** San Francisco, CA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Systems Engineering, Cloud Computing, Computer Clusters, Computer Engineering, System Configuration, Data Centers, Linux, File Systems, Distributed Computing Environment, Network Topologies, InfiniBand, Open Source Technology, Remote Direct Memory Access, AI Infrastructure, Ceph (Software), Pytorch, Large Language Models, Multi-Agent Systems, HybridCloud, Kubernetes, Information Technology, Slurm, TensorRT - **Published:** August 4, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=1d46d5bdc44a2d17 ## About the Role + Experience: 8-10+ years as a Solution Architect, Systems Engineer, or Infrastructure Architect handling enterprise cloud, HPC, or AI environments. + AI & Hardware Mastery: Deep hands-on experience with high-density GPU nodes (NVIDIA HGX/DGX/Blackwell), Linux bare-metal virtualization, and RDMA networking (InfiniBand/RoCE v2). + Orchestration & Storage: Proficiency with Kubernetes, container networking (CNI), and parallel file systems (Lustre, VAST, Ceph). + Customer-Facing Execution: Proven ability to lead complex technical workshops, present to C-suite executives, and resolve high-stakes production bottlenecks. * Preferred Qualifications: + Practical experience tuning NCCL communication, GPUDirect Storage (GDS), or custom InfiniBand routing fabrics. + Contributions to open-source cloud-native or AI orchestration projects (e.g., CNCF, Ray, Slurm, k0rdent). + B.S. or M.S. in Computer Science, Computer Engineering, or quantitative discipline. ## Description With 100% renewable energy, we build, own and operate our data centers and take pride in being at the forefront of sustainable solutions for the ever-evolving applications of high-performance compute. We believe that human progress is invaluable, but it should be done in the right way - responsibly, sustainably and having a positive impact on the communities we operate in. As a Principal Solution Architect, you will sit at the tip of the spear for IREN's AI Cloud platform. You will partner directly with client CTOs, lead ML engineers, and infrastructure architects to design, benchmark, and optimize massively distributed training and inference clusters. You will convert complex customer requirements into optimized hardware/software configurations while creating reference architectures that set industry benchmarks., * Primary Technical Authority: Serve as the trusted lead technical architect for tier-1 enterprise customers migrating or scaling AI workloads on IREN's infrastructure. * Architectural Blueprinting: Design, document, and publish end-to-end reference architectures covering GPU clusters, InfiniBand/RoCE network topologies, Ceph/Lustre storage fabrics, and Kubernetes multi-tenancy. * PoCs & Performance Benchmarking: Lead hands-on proof-of-concept deployments, conducting cluster benchmarking (e.g., NCCL-tests, Linpack, MLPerf, IO500) to validate customer SLA requirements. * Workload Optimization: Partner with client engineering teams to profile and optimize distributed training jobs (Megatron-LM, DeepSpeed, PyTorch FSDP) and inference pipelines (vLLM, TensorRT-LLM). * Product Feedback Conduit: Translate customer pain points, telemetry, and edge cases into structured platform requirements for IREN's core Product and Operations teams. ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Reference Architecture of AI in the Cloud](https://www.wearedevelopers.com/videos/1613-reference-architecture-of-ai-in-the-cloud) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Serverless deployment of (large) NLP models ](https://www.wearedevelopers.com/videos/158-serverless-deployment-of-large-nlp-models) ## Related Articles - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift)