> Markdown version of [/jobs/ext/2722290-member-of-technical-staff-ai-cloud-infrastructure](https://www.wearedevelopers.com/jobs/ext/2722290-member-of-technical-staff-ai-cloud-infrastructure). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Member of Technical Staff - AI Cloud Infrastructure - **Company:** EMERALD AI LLC - **Location:** Oakland, United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Amazon S3, Cloud Computing, Software Debugging, Linux, General Parallel File Systems, InfiniBand, Python (Programming Language), Kerberos (Protocol), Machine Learning, Network Control, Performance Tuning, Remote Direct Memory Access, Cloud Services, Ansible, Systems Integration, Virtual Local Area Networks, Weka, Ceph (Software), AI Platforms, Kubernetes, Slurm, Hardware Infrastructure, Software Coding, Terraform - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/member-of-technical-staff-ai-cloud-infrastructure-emerald-ai-8935887 ## About the Role * At least 7+ years of experience in infrastructure or platform engineering, including the architecture and launch of a managed cloud or AI platform that reached production users. * Strong experience with Kubernetes and Slurm and offering them as managed service * Production experience deploying or operating Lustre or a comparable parallel filesystem such as GPFS, Weka, VAST, or BeeGFS, with a solid understanding of parallel filesystem architecture, tuning, and failure modes. * A strong grasp of cloud service fundamentals, including control planes, tenancy and isolation models, APIs, quota and metering systems, and the operational discipline of running a service that customers pay for. * Deep Linux systems knowledge, mature infrastructure as code practice with tools such as Terraform and Ansible, and solid programming ability in Python or Go. * Familiarity with GPU infrastructure, including high performance networking with InfiniBand, RoCE, and RDMA, and the GPU software stack. Preferred requirements * Prior time at a GPU cloud, a hyperscaler AI service, or an HPC center that delivers compute and storage as a service, especially one built on rented or colocated capacity. * Familiarity with NVIDIA reference architectures such as SuperPOD, along with GPUDirect Storage, NCCL debugging, and DCGM. * Experience with Lustre multitenancy features such as nodemap, fileset mounts, and Kerberos, or with service provider deployments of VAST or Weka. * Experience negotiating with and integrating multiple infrastructure vendors, together with a practice of designing for portability between them. * Experience running object storage at scale with systems such as S3, Ceph, or MinIO, including the design of data tiering. * Experience building billing, metering, or FinOps pipelines for services that charge by usage. ## Description Emerald AI is building the world's first power flexible managed cloud infrastructure. We are hiring a senior infrastructure engineer to architect and stand up our managed cloud services from end to end. The work covers the platform, the control plane, and the customer experience that together make up a managed AI cloud. The right person has done this before. They have built or served as a core early engineer on a managed cloud or AI platform, whether at a GPU cloud, an internal machine learning platform run at scale, a hyperscaler AI service, or a HPC research computing center operated as a service. This is a role for an architect who still builds. You will make the major design decisions and then implement them yourself., * Architect our managed services from 0*1. Define the productization of GPU capacity, encompassing isolation boundaries, tenant models, provisioning flows, and service catalogs that scale across diverse providers. * Engineer the platform core. Build robust control-plane services, self-service customer interfaces, and automated lifecycle systems, including usage metering integrated with billing infrastructure. * Onboard and vet infrastructure partners. Conduct deep technical assessments of bare-metal GPU vendors, evaluating fabric quality, network isolation, and economics to automate the path from handoff to active tenant. * Design end-to-end multi-tenancy. Implement rigorous isolation across compute, storage, and networking (InfiniBand/VLANs), ensuring secure boundaries, QoS, and encryption even when customers possess root access. * Drive workload orchestration. Manage Kubernetes and Slurm environments for large-scale training and inference, overseeing node health, driver fleets, and kernel management across heterogeneous clouds. * Lead high-performance storage strategy. Deploy and integrate parallel storage solutions like Lustre, VAST, or Weka, leveraging your deep experience with these systems to ensure they fold cleanly into our provisioning model. * Ensure operational excellence. Define SLOs, observability standards, and incident response protocols that bridge our internal standards with underlying provider SLAs to deliver a reliable, sellable product. ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) - [Discover the open source trio you didn’t expect: .NET and PostgreSQL on Linux](https://www.wearedevelopers.com/videos/2042-discover-the-open-source-trio-you-didn-t-expect-net-and-postgresql-on-linux) - [AI Factories at Scale](https://www.wearedevelopers.com/videos/1139-ai-factories-at-scale) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development)