> Markdown version of [/jobs/ext/2283314-senior-engineering-manager-object-storage-dgx-cloud](https://www.wearedevelopers.com/jobs/ext/2283314-senior-engineering-manager-object-storage-dgx-cloud). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Engineering Manager, Object Storage - DGX Cloud - **Company:** NVIDIA Corporation - **Location:** Santa Clara, CA, United States - **Experience:** Expert - **Salary:** $272,000.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Amazon S3, Automation of Tests, C++ (Programming Language), Cloud Computing, Cloud Storage, Software Quality, Continuous Integration, Corona (Software Development Kit), Data Files, Extract Transform Load (ETL), Distributed Data Store, Infrastructure as a Service (IaaS), Python (Programming Language), Platform as a Service (PAAS), Software Engineering, Management of Software Versions, AI Infrastructure, Information Technology, Data Pipelines, Golang - **Published:** August 28, 2026 - **Apply:** https://nvidia.wd5.myworkdayjobs.com/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Senior-Engineering-Manager--Object-Storage---DGX-Cloud_JR2021079-1 ## About the Role * BS, MS, or PhD in Computer Science, Electrical Engineering, or a related field - or equivalent experience. * 10+ overall years of software engineering experience, including 4+ years in an engineering management role leading teams of 10 or more engineers delivering production services at scale. * Deep technical background in distributed storage systems, object storage platforms, or large-scale cloud data services; hands-on development experience in Go, C++, Python, or equivalent systems languages. * Direct, hands-on experience building or scaling S3-compatible object storage systems in a production cloud or private cloud environment - with demonstrable improvements in throughput, durability, or operational efficiency. * Demonstrated experience building or operating cloud storage services - with accountability for reliability, performance, and capacity at scale in a production environment. * Proven track record of shipping production software on time - managing scope, risk, and delivery across multiple concurrent workstreams. * Strong experience with modern software development and service delivery practices: CI/CD, automated testing, SLO-based reliability, production observability, and incident management. * Demonstrated ability to attract, develop, and retain strong engineering talent in a driven environment, with a track record of growing engineers into senior and staff-level roles. * Excellent written and verbal communication - able to translate complex technical trade-offs for product partners and engineering constraints for executive audiences. Ways to Stand Out from the crowd: * Prior experience designing and operating internal cloud storage services (IaaS/PaaS) with well-defined SLAs, metered usage, and internal customer-facing APIs. * Background in data movement, data staging, or prefetching tooling for AI/ML workloads - with direct experience optimizing data pipelines to reduce GPU idle time during training or inference. * Familiarity with AI infrastructure storage patterns: checkpoint storage, dataset versioning, write-once-read-many (WORM) access patterns, or storage-aware scheduling at 10k+ GPU scale. Experience managing capacity planning, cost optimization, and chargeback modeling for shared internal storage infrastructure. * Track record of adopting AI-assisted development tools to meaningfully improve team productivity, with concrete examples. History of growing engineers into senior ICs or leads, and building diverse, inclusive teams with strong retention. NVIDIA's Object Storage Platform and data movement tooling form a critical layer in keeping NVIDIA's GPU fleet productive - every model trained, every checkpoint saved, and every dataset staged passes through the systems this team builds and operates. This is a high-impact role at the center of NVIDIA's AI infrastructure. ## Description * Own roadmap execution for NVIDIA's internal object storage service - partnering with internal customers, Product Management, and Architecture to translate multi-quarter goals into clear engineering plans with measurable milestones. * Drive development and operation of NVIDIA's S3-compatible object storage service, ensuring it meets the performance, durability, availability, and scalability demands of AI training and inference workloads at exabyte scale. * Lead the Data Movement Tools team in building and evolving tooling that stages datasets, model checkpoints, and artifacts from distributed storage to GPU-adjacent compute - minimizing I/O bottlenecks and keeping accelerators fully utilized. * Define and uphold service reliability standards: SLOs, capacity planning, incident response, root cause analysis, and on-call hygiene. Partner with SRE to ensure the platform meets the availability commitments internal customers depend on. * Establish and enforce engineering standards across both teams: design reviews, code quality, CI/CD practices, automated testing, and production observability. Recruit, mentor, and develop engineers across all levels, conducting regular 1:1s, performance cycles, and career growth conversations. Build a diverse, inclusive, and high-retention team. * Collaborate closely with SRE, Platform, Networking, and Security teams to ensure smooth transitions from development to production and rapid resolution of customer-impacting issues. * Champion the adoption of AI-assisted development tooling - coding assistants, agentic workflows, and automated testing harnesses - to accelerate team productivity and raise engineering output. Represent the Object Storage engineering organization to senior leadership, providing transparent status updates, surfacing risks early, and advocating for the resources needed to succeed. ## Related Videos - [Your Next AI Needs 10,000 GPUs. Now What?](https://www.wearedevelopers.com/videos/1590-your-next-ai-needs-10-000-gpus-now-what) - [Introduction to TXT](https://www.wearedevelopers.com/videos/30-introduction-to-txt) - [WeAreDevelopers LIVE - CSS is DOOMed](https://www.wearedevelopers.com/videos/1838-wearedevelopers-live-css-is-doomed) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Implementing continuous delivery in a data processing pipeline](https://www.wearedevelopers.com/videos/73-implementing-continuous-delivery-in-a-data-processing-pipeline) - [Hate organising your photos? Try it with 5 Terabytes](https://www.wearedevelopers.com/videos/79-hate-organising-your-photos-try-it-with-5-terabytes) ## Related Articles - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers)