> Markdown version of [/jobs/ext/1323552-member-of-technical-staff-engineering-lead-compute-platform](https://www.wearedevelopers.com/jobs/ext/1323552-member-of-technical-staff-engineering-lead-compute-platform). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Member of Technical Staff - Engineering Lead, Compute Platform - **Company:** REFLECTION LLC - **Location:** United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Computing Platforms, Systems Engineering, Cloud Storage, Data Centers, Software Debugging, Fault Tolerance, Software Architecture, Graphics Processing Unit (GPU), Multi-Cloud, Kubernetes, Database Replication, Hardware Debugging - **Published:** July 17, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=8798facb90cbad16 ## About the Role * Experience building, mentoring, and growing systems or infrastructure teams while staying technically hands-on. (Comfortable growing into leading a team of ~10 quickly if you haven't managed at that scale before.) * Deep systems-level engineering experience with a focus on cluster-wide behavior and maintenance. * Strong coding ability and the credibility to earn the technical trust of a strong team. * Depth in at least one of orchestration, storage, or GPU hardware - with the ability to learn the rest. Deep GPU knowledge beyond standard Kubernetes (e.g., NCCL) is a plus, not a prerequisite. * Alignment with a Kubernetes-first architecture. * Cloud storage expertise - managing high-performance data products (like VAST) across multiple data centers and handling datasets and checkpointing at scale - is a plus. * Experience managing vendors, including negotiating and operating important deals. * Ability to guide strategy and drive execution across a multi-cloud, large-fleet environment, and to partner effectively with research and training teams. ## Description Reflection's Compute Platform team keeps our compute layer healthy and highly available. We run a Kubernetes-based platform distributed across multiple neo-clouds, where multi-cloud scheduling, node health, and performance debugging at scale present genuinely hard systems problems. As Compute Platform Lead, you'll provide front-line leadership of the team that builds and operates this layer. You'll build, mentor, and grow a team of strong systems engineers, guide the technical and architectural decisions across multi-cloud scheduling, cluster management, and next-generation GPU deployments, and work closely with our training teams to co-design fault tolerance, node health checks, and remediation. You'll stay close enough to the systems to make targeted contributions as an individual contributor and to maintain a deep understanding of the compute fleet our largest training runs depend on. Managing vendors - and the important deals that come with them - is a core part of the job., * Build, mentor, and grow a high-performing team of systems engineers. Coach and support your reports in understanding, and pursuing, their professional growth. * Provide front-line leadership of engineering efforts to keep the compute fleet reliable and highly available - multi-cloud scheduling, cluster management, and the path to next-generation GPUs and increasingly larger cluster sizes. * Stay hands-on: become familiar with the team's technical stack enough to make targeted contributions as an individual contributor. * Manage day-to-day execution: prioritize the team's work and manage projects in a highly dynamic, fast-paced environment. * Guide technical and architectural decisions, emphasizing scalability, robustness, and reliability - automatic remediation, topology-aware scheduling, capacity planning, rapid hardware debugging, and cluster-wide monitoring and performance benchmarking. * Work closely with our training teams to co-design fault tolerance, node health checks, and remediation, and manage the vendor relationships and important deals the compute fleet depends on. * Prepare the fleet for what's next: next-generation GPUs and larger clusters, and - longer term - multi-cloud storage, petabyte-scale data replication, and GPU-to-GPU network performance. * Raise the bar for technical judgment, prioritization, communication, and execution in a fast-moving environment. ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [The Open-source Java SDK for Multi-Cloud Development - Sandeep Pal](https://www.wearedevelopers.com/videos/2113-the-open-source-java-sdk-for-multi-cloud-development-sandeep-pal) - [The Sustainability Race: AI's Promises, Pitfalls and Potential](https://www.wearedevelopers.com/videos/100155-the-sustainability-race-ai-s-promises-pitfalls-and-potential) - [A Deep Dive on How To Leverage the NVIDIA GB200 for Ultra-Fast Training and Inference on Kubernetes](https://www.wearedevelopers.com/videos/1625-a-deep-dive-on-how-to-leverage-the-nvidia-gb200-for-ultra-fast-training-and-inference-on-kubernetes) - [Instant KAI Sandboxes with vCluster: Multi-Tenant, Multi-Scheduler GPU Sharing](https://www.wearedevelopers.com/videos/100333-instant-kai-sandboxes-with-vcluster-multi-tenant-multi-scheduler-gpu-sharing) ## Related Articles - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Learning Kubernetes made easy with KubeCampus](https://www.wearedevelopers.com/magazine/348-learning-kubernetes-made-easy-with-kubecampus) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [What Makes WeAreDevelopers World Congress Different From Every Other Tech Event?](https://www.wearedevelopers.com/magazine/701-what-makes-wearedevelopers-world-congress-different-from-every-other-tech-event)