> Markdown version of [/jobs/ext/2989754-cloud-hardware-storage-engineer](https://www.wearedevelopers.com/jobs/ext/2989754-cloud-hardware-storage-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Cloud Hardware Storage Engineer - **Company:** Microsoft - **Location:** San Francisco, CA, United States - **Experience:** Expert - **Salary:** $119,800.0 - $234,700.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Microsoft Azure, Microsoft Online Services, C Sharp (Programming Language), C++ (Programming Language), Computer Engineering, Continuous Integration, Firmware, Python (Programming Language), PCI Express, Kusto Query Language, SQL Databases, Management of Software Versions, Storage Devices, Large Language Models, Information Technology, Data Management, Machine Learning Operations, Hardware Infrastructure, Data Pipelines, Nvme - **Published:** September 18, 2026 - **Apply:** https://www.dice.com/job-detail/8b19fc5a-cb21-44c6-ae17-e5a11f6235f6 ## About the Role Master's Degree in Electrical Engineering, Computer Engineering, Mechanical Engineering, or related field AND 3+ years technical engineering experience OR Bachelor's Degree in Electrical Engineering, Computer Engineering, Mechanical Engineering, or related field AND 5+ years technical engineering experience OR equivalent experience, * Ability to meet Microsoft, customer and/or government security screening requirements are required for this role. These requirements include but are not limited to the following specialized security screenings: * Microsoft Cloud Background Check: This position will be required to pass the Microsoft Cloud Background Check upon hire/transfer and every two years thereafter., o Bachelor's Degree in Computer Engineering, Computer Science, Electrical Engineering, or related field AND 8+ years of firmware or embedded systems engineering experience OR Master's Degree in Computer Engineering, Computer Science, Electrical Engineering, or related field AND 6+ years of firmware or embedded systems engineering experience OR equivalent experience # 6+ years developing SSD or storage device firmware, including 4+ years working directly with NVMe and PCIe protocols # Demonstrated depth in storage device resiliency and fault analysis - failure mode characterization, error handling and recovery paths, and root-cause investigation of field failures # Experience supporting live-site operations for storage at fleet scale, including on-call ownership and production incident resolution # Track record of owning end-to-end technical design across the full reliability lifecycle: detection, prediction, mitigation, and repair # Proven experience building automation-heavy systems that operate safely at hyperscale, with the guardrails, staged rollout, and blast-radius controls that requires Hardware Engineering IC4 - The typical base pay range for this role across the U.S. is USD $119,800 - $234,700 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $160,200 - $261,000 per year. ## Description What you'll do * Build the platform behind Azure's fault self-healing and failure-prediction system with telemetry pipelines, prediction services, decision logic, and automated repair workflows. * Ship AI agents to production: prompt and tool design, retrieval over diagnostics data, evaluation harnesses, guardrails, and the CI/CD path that deploys new agent skills safely. * Close the loop safely: automated remediation with staged rollout, blast-radius limits, and verification. * Own the developer experience: SDKs, APIs, and dashboards that let engineers across the org author and deploy new detection and repair logic themselves. * Build and monitor measurements: prediction precision/recall, false-repair rate, action success rate, and regression gates that block a bad model from shipping. What you'll bring * 8+ years building and operating production software, distributed services, data platforms, or large-scale automation. * Python and C# (or C++/Rust), plus cloud-scale data pipelines at high volume. * AI/ML in production, not just experimentation: agent or model serving, evaluation, versioning and rollback, drift and regression monitoring. * LLM application patterns: agent/tool-calling, RAG, structured output and judgment about when an LLM is the wrong answer. * A track record of automation that takes real actions on real infrastructure, with the safety engineering that requires. * Bonus: anomaly detection on time-series data; Kusto/SQL; server hardware, firmware, or datacenter operations. Why this role * Your impact shows up in fleet availability metrics and customer uptime, quarter over quarter. The AI agents you ship take autonomous action on production infrastructure at a scale very few teams can offer. ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Playing Pong on a shoulder press machine](https://www.wearedevelopers.com/videos/100140-playing-pong-on-a-shoulder-press-machine) - [Developer Tools for Microsoft Azure](https://www.wearedevelopers.com/videos/450-developer-tools-for-microsoft-azure) - [Agent Smith Gets Hardware: Autonomous IoT Hacking From Debug Port to Cloud API](https://www.wearedevelopers.com/videos/100258-agent-smith-gets-hardware-autonomous-iot-hacking-from-debug-port-to-cloud-api) - [Azure-Well Architected Framework - designing mission critical workloads in practice](https://www.wearedevelopers.com/videos/1529-azure-well-architected-framework-designing-mission-critical-workloads-in-practice) - [When Agents Meet Legacy: Never Change a Running System](https://www.wearedevelopers.com/videos/100320-when-agents-meet-legacy-never-change-a-running-system) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Top-Paying Tech Jobs (with Salaries)](https://www.wearedevelopers.com/magazine/372-top-paying-tech-jobs-with-salaries) - [Software Engineer Salary London](https://www.wearedevelopers.com/magazine/252-software-engineer-salary-london) - [How Much Does a Software Engineer Make? Realistic Software Engineering Salaries](https://www.wearedevelopers.com/magazine/425-how-much-does-a-software-engineer-make-realistic-software-engineering-salaries) - [Highest Paying Tech Companies in Europe](https://www.wearedevelopers.com/magazine/162-highest-paying-tech-companies-in-europe) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk)