> Markdown version of [/jobs/ext/3252296-software-engineer-platform](https://www.wearedevelopers.com/jobs/ext/3252296-software-engineer-platform). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer, Platform - **Company:** fal - Features & Labels - **Location:** San Francisco, CA, United States - **Experience:** Experienced - **Salary:** $180,000.0 - $250,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Border Gateway Protocol, Booting (BIOS), Cloud Computing, Configuration Management, Profiling, Nvidia CUDA, Linux, RAID, File Systems, Network Interface Controllers, General Parallel File Systems, InfiniBand, Python (Programming Language), Key Management, Network Troubleshooting, Logical Volume Manager, Network Configuration and Change Management, Network File Systems, Performance Tuning, Remote Direct Memory Access, Ansible, Software Engineering, Tcpdump, Virtual Local Area Networks, Graphics Processing Unit (GPU), Build Server, Selinux, Storage Technologies, Bare Metal, Build Tools, Hardware Infrastructure, Terraform, Nvme, Vulnerability Analysis - **Published:** September 20, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=f23613f8d19ae320 ## About the Role * 3+ years experience managing bare-metal and cloud based server fleets at scale (100+ nodes) * Strong software engineering skills in Python; you write production tooling, not scripts * Deep Linux systems knowledge: boot process, kernel tuning, networking, storage, systemd, cgroups, namespaces, performance profiling * Strong experience with configuration management and infrastructure-as-code: Ansible, Terraform, cloud-init * Solid understanding of storage technologies: LVM, RAID, NVMe, NFS, Lustre or GPFS, and Linux I/O stack tuning * Familiarity with hardware diagnostics and failure modes (GPUs, NVMe, NICs, memory) * Experience building internal tools or dashboards for infrastructure visibility * Excellent communication and ability to drive technical decisions across teams * Self-starter who executes quickly, takes ownership, and constantly seeks improvement, * Familiarity with network configuration and diagnostics (VLAN, VXLAN, ECMP, BGP, tcpdump) * Experience with NVIDIA GPU infrastructure: driver management, health monitoring, DCGM, NVLink/NVSwitch diagnostics, RDMA, InfiniBand/RoCEv2 * Experience with AMD GPUs * Experience with bare metal and VM provisioning (PXE/iPXE, Kickstart, libvirt, Qemu/KVM) * Experience with compliance frameworks relevant to cloud providers (SOC 2, ISO 27001) ## Description * Build and maintain Python fleet tracking system that manages the full lifecycle of servers including contracting and procurement, target use, pricing, availability, health, RMAs, etc * Build server management tooling that automates provisioning, health checks, GPU diagnostics, recovery and alerting * Create and maintain metrics, dashboards, and alerting for hardware health across the fleet (GPU errors, disk failures, network issues, thermals) * Leverage AI to an extreme level to build tools and automate alerting and recovery * Implement and enforce OS-level security: hardening baselines, SELinux/AppArmor policies, SSH key management, vulnerability scanning, and compliance automation * Manage and optimize distributed and local storage systems supporting model weights, checkpoints, and ephemeral scratch: NVMe arrays, NFS, parallel file systems, and object storage * Tune Linux systems for AI workloads: kernel parameters, NUMA topology, CPU pinning, hugepages, I/O schedulers, and GPU driver stack optimization (NVIDIA drivers, CUDA, container runtimes) * Develop a suite of automated error detection and recovery processes * Work with partners to solve technical issues ## Related Videos - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) - [Reducing Cognitive Overload Through Platform Engineering](https://www.wearedevelopers.com/videos/679-reducing-cognitive-overload-through-platform-engineering) - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Discover the open source trio you didn’t expect: .NET and PostgreSQL on Linux](https://www.wearedevelopers.com/videos/2042-discover-the-open-source-trio-you-didn-t-expect-net-and-postgresql-on-linux) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Dev Digest 129 - Now that's what I call private data!](https://www.wearedevelopers.com/magazine/468-dev-digest-129-now-that-s-what-i-call-private-data) - [The Best X (Twitter) Accounts for Developers](https://www.wearedevelopers.com/magazine/294-the-best-x-twitter-accounts-for-developers)