Senior Software Engineering Manager - KV Cache Platform
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+10 more
Job description
DDN is seeking a Senior Software Engineering Manager to lead the engineering organization responsible for our KV Cache Platform-a distributed memory and storage platform that accelerates large-scale LLM inference across GPU clusters.
In this role, you will lead geographically distributed engineering teams responsible for building highly scalable, low-latency distributed systems that power AI inference. You will define the technical vision and execution strategy for the platform while partnering closely with Product Management, Sales, Customer Engineering, NVIDIA, and executive leadership to deliver innovative AI infrastructure that meets customer needs and supports DDNâs long-term product strategy.
This is a highly visible leadership role with responsibility for engineering execution, customer success, roadmap delivery, and building a world-class engineering organization.
Responsibilities Lead, mentor, and grow a geographically distributed team of software engineers and technical leaders, fostering a culture of technical excellence, innovation, ownership, and collaboration.
- Define and execute the technical strategy and roadmap for the KV Cache Platform, ensuring scalability, reliability, security, and operational excellence.
- Drive the architecture, development, and delivery of distributed systems supporting AI inference, GPU memory optimization, distributed caching, RDMA networking, GPUDirect Storage, NVIDIA BlueField DPUs, and emerging AI infrastructure technologies.
- Partner closely with Product Management, Sales, Customer Engineering, NVIDIA, and strategic technology partners to prioritize customer requirements, drive proof-of-concepts (POCs), influence product direction, and successfully deliver customer deployments.
- Own day-to-day engineering execution, including feature development, release planning, bug triage, production issues, customer escalations, and cross-functional execution to ensure timely, high-quality software delivery.
- Establish engineering best practices for software quality, observability, automation, performance, testing, and production readiness.
- Collaborate across engineering, infrastructure, and hardware teams to deliver scalable, production-ready AI infrastructure while developing future engineering leaders and driving continuous improvement.
Requirements
- 15+ years of experience building distributed systems, cloud infrastructure, storage platforms, or AI infrastructure software.
- 7+ years leading high-performing software engineering organizations, including geographically distributed teams.
- Proven experience delivering large-scale distributed infrastructure products from architecture through production deployment.
- Strong background in distributed systems, Linux, networking, performance engineering, and cloud-native architectures.
- Hands-on programming experience with Go and Python; experience with C/C++ is a plus.
- Demonstrated ability to lead cross-functional initiatives and influence technical direction across multiple organizations.
- Experience building AI infrastructure, LLM serving platforms, distributed caching systems, or high-performance storage solutions.
- Experience with technologies such as NVIDIA Dynamo, TensorRT-LLM, Triton, RDMA, GPUDirect Storage, BlueField DPUs, Kubernetes, or related AI infrastructure.
- Background in HPC, distributed storage, networking, or enterprise infrastructure software.
- Experience working directly with strategic customers, technology partners, OEMs, or hyperscalers to deliver enterprise AI solutions.
About the company
This is an incredible opportunity to be part of a company that has been at the forefront of AI and high-performance data storage innovation for over two decades. DataDirect Networks (DDN) is a global market leader renowned for powering many of the worldâs most demanding AI data centers, in industries ranging from life sciences and healthcare to financial services, autonomous cars, Government, academia, research and manufacturing.
âDDNâs A3I solutions are transforming the landscape of AI infrastructure.â - IDC
| âThe real differentiator is DDN. I never hesitate to recommend DDN. DDN is the de facto name for AI Storage in high performance environmentsâ - Marc Hamilton, VP, Solutions Architecture & Engineering | NVIDIA |
DDN is the global leader in AI and multi-cloud data management at scale. Our cutting-edge data intelligence platform is designed to accelerate AI workloads, enabling organizations to extract maximum value from their data. With a proven track record of performance, reliability, and scalability, DDN empowers businesses to tackle the most challenging AI and data-intensive workloads with confidence.
Our success is driven by our unwavering commitment to innovation, customer-centricity, and a team of passionate professionals who bring their expertise and dedication to every project. This is a chance to make a significant impact at a company that is shaping the future of AI and data management.
Our commitment to innovation, customer success, and market leadership makes this an exciting and rewarding role for a driven professional looking to make a lasting impact in the world of AI and data storage., Why Join DDN?
At DDN, youâll lead one of our most strategic AI infrastructure initiatives and help define the future of distributed memory systems for large-scale AI inference. Youâll work alongside world-class engineers and strategic partners, solve some of the industryâs most challenging distributed systems problems, and build technology that powers the next generation of AI deployments worldwide.
Salary Range for this role: $220,000 - $275,000 DDN
Join our dynamic and driven team, where engineering excellence is at the heart of everything we do. We seek individuals who love to challenge themselves and are fueled by curiosity. Here, youâll have the opportunity to work across various areas of the company, thanks to our flat organizational structure that encourages hands-on involvement and direct contributions to our mission. Leadership is earned by those who take initiative and consistently deliver outstanding results, both in their work ethic and deliverables, making strong prioritization skills essential. Additionally, we value strong communication skills in all our engineers and researchers, as they are crucial for the success of our teams and the company as a whole.
Interview Process: After submitting your application, one of our recruiters will review your resume. If your application passes this stage, you will be invited to a 30-minute interview during which a member of our team will ask some basic questions. If you clear the interview, you will enter the main process, which can consist of up to four interviews in total:
- Coding assessment: Often in a language of your choice.
- Systems design: Translate high-level requirements into a scalable, fault-tolerant service (depending on role).
- Real-time problem-solving: Demonstrate practical skills in a live problem-solving session.
- Meet and greet with the wider team.
- Our goal is to finish the main process in 2-3 weeks at most.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on diversityjobs.comGood distractions
Talks and stories from around this role â technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence
Dev Digest 120 - Apple and peers
Dev Digest 121 - AI goes offline
Dev Digest 132 - Binging WADFlix?