> Markdown version of [/jobs/ext/2702848-software-engineer-infrastructure-storage](https://www.wearedevelopers.com/jobs/ext/2702848-software-engineer-infrastructure-storage). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer - Infrastructure Storage - **Company:** Lambda Inc. - **Location:** San Francisco, CA, United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Amazon S3, Computer Programming, Data Centers, Software Debugging, File Systems, Distributed Data Store, Distributed Systems, Internet Small Computer System Interface (ISCSI), NetApp Applications, Operational Databases, Performance Tuning, Software Product Management, Remote Direct Memory Access, Scientific Computating, Software Engineering, Weka, AI Infrastructure, Ceph (Software), Storage Technologies, Block Storage, Nvme - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/senior-software-engineer-infrastructure-storage-lambda-3-8411128 ## About the Role + 10+ years of experience in storage engineering with at least 5+ years in a management or lead role. + Systems-Level Programming and Architecture + Storage Protocol and API Mastery: + Storage Performance Optimization + DPKD SPKD + Physical Infrastructure Knowledge + Operational Acumen * Technical Skills: + Experience in serving one or more of the following storage protocols: object storage (e.g., S3), block storage (e.g., iSCSI), or file storage (e.g., NFS, SMB, Lustre). + Professional individual contributor experience as a storage engineer or storage SRE. + Familiarity with modern storage technologies (e.g., NVMe, RDMA, DPUs) and their role in optimizing performance. * People Management: + Experience building a high-performance team through deliberate hiring, upskilling, planned skills redundancy, performance-management, and expectation setting., + Experience driving cross-functional engineering management initiatives (coordinating events, strategic planning, coordinating large projects). + Experience with NVidia SuperNIC DPUs for edge-caching (such as implementing GPUDirect Storage). * Technical Skills: + Deep experience with Vast, Weka and/or NetApp in an HPC or AI Infrastructure environment. + Deep experience implementing CEPH in an HPC or AI infrastructure environment at a scale greater than 100PB. * People Management: + Experience driving organizational improvements (processes, systems, etc.) + Experience training, or managing managers. ## Description We are seeking a seasoned Storage Software Engineer with experience designing and deploying various storage protocol solutions at scale (object, block, and file). This is a unique opportunity to work at the intersection of large-scale distributed systems and the rapidly evolving field of artificial intelligence infrastructure. This is an opportunity to have a significant impact on the future of AI. You will be building the foundational infrastructure that powers some of the most advanced AI research and products in the world. What You'll Do * Technical Leadership: + At the Senior Level * Execution: + Systems-Level Programming and Architecture + Design, develop, and maintain software for storage systems, focusing on performance, scalability, and reliability. + Implement and optimize storage protocol APIs for file (e.g., NFS, SMB), block (e.g., Fibre Channel), and object (e.g., S3) access. + Develop distributed systems for managing and orchestrating storage resources across multiple storage solutions and redundant arrays. + Collaborate with hardware and system architects to integrate software with various storage solutions, including NVMe and GPU-direct storage. + Troubleshoot and debug complex issues in a production data center environment. + Contribute to the full software development lifecycle, from requirements gathering and design to deployment and maintenance. * Collaboration + Work closely with the storage software teams and networking teams to execute on cross-functional infrastructure initiatives and new data-center deployments including integration of storage protocols across a variety of on-prem storage solutions. + Work closely with the control plane and MK8s teams to meet customer/product requirements for usability, reliability, and telemetry. + Work with the observability team to build/track SLOs/SLIs. + Work closely with Networking, Compute, and Storage Software Engineering teams to deploy high-performance distributed storage solutions to serve AI/ML workloads. + Partner with the fleet engineering team to ensure seamless deployment, monitoring, and maintenance of the distributed storage solutions. * Innovate: + Stay current with the latest trends and research into AI and HPC storage technologies. + Work with the Lambda product team to uncover new trends in the AI inference and training product category that will inform emerging storage solutions. + Optimize protocol solutions for the AI product vertical exploring optimizations for AI Inference, training, and scientific computing applications. ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [WeAreDevelopers LIVE - CSS is DOOMed](https://www.wearedevelopers.com/videos/1838-wearedevelopers-live-css-is-doomed) - [It's all about the Data](https://www.wearedevelopers.com/videos/425-it-s-all-about-the-data) - [Hate organising your photos? Try it with 5 Terabytes](https://www.wearedevelopers.com/videos/79-hate-organising-your-photos-try-it-with-5-terabytes) - [Building Systems that Last](https://www.wearedevelopers.com/videos/1389-building-systems-that-last) - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [What does the history of data storage tell us about the future?](https://www.wearedevelopers.com/magazine/495-what-does-the-history-of-data-storage-tell-us-about-the-future) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)