Senior Cloud Hardware Storage Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+7 more
Job description
Microsoft Silicon and Cloud Hardware Infrastructure Engineering (SCHIE) is the team behind Microsoftâs expanding Cloud Infrastructure and responsible for powering Microsoftâs âIntelligent Cloudâ mission. CHIE delivers the core infrastructure and foundational technologies for Microsoftâs over 200 online businesses including Bing, MSN, Office 365, Xbox Live, Skype, OneDrive and the Microsoft Azure platform globally with our server and data center infrastructure, security and compliance, operations, globalization, and manageability solutions. Our focus is on smart growth, high efficiency, and delivering a trusted experience to customers and partners worldwide and we are looking for passionate, high-energy engineers to help achieve that mission.
As Microsoftâs cloud business continues to grow the ability to deploy new offerings and HW infrastructure on time, in high volume with high quality and lowest cost is of paramount importance. To achieve this goal, the Silicon Cloud Hardware Infrastructure Engineering (SCHIE) team is instrumental in defining and delivering measures of success for hardware design, qualification, fleet support, scale, and sustainability related to Microsoft cloud hardware.
Azure Memory and Storage Center of Excellence (AMS CoE) is part of the SCHIE organization focusing on Memory and Storage devices going into the Cloud hardware servers. AMS provide memory and storage solutions to Azure, drive memory and storage suppliers to deliver high quality products, meeting our requirements.
We are looking for a Senior Cloud Hardware Engineer to scale Azureâs Fault Self-Healing and Failure Prediction systems.
You will own the end-to-end technical design and execution of the fault prevention ecosystem, spanning telemetry, automation, isolation logic, firmware interactions, and repair workflows, operating at hyperscale across millions of nodes. The role directly impacts customer uptime and fleet availability.
Responsibilities
-
Design and build best-in-class fleet resiliency systems for storage devices at scale
-
Develop scalable live monitoring capabilities, fault detection and repair solutions
-
Design features for SSDs, HDDs and Storage Accelerator firmware deployment
-
Lead collaboration projects with hardware, firmware and software teams that fault reduction projects
-
Build automation to drive repair efficiency for storage operations in the production fleet
-
Collaborate with suppliers to design reliable, high performance and quality storage devices
-
Analyze data to identify, prototype, and drive the implementation of technical and process improvements to increase the predictability, agility, and quality of Azure systems
-
Actively support Azure service stakeholders
Requirements
- Masterâs Degree in Electrical Engineering, Computer Engineering, Mechanical Engineering, or related field AND 3+ years technical engineering experience OR Bachelorâs Degree in Electrical Engineering, Computer Engineering, Mechanical Engineering, or related field AND 5+ years technical engineering experience OR equivalent experience
Other Requirements:
- Ability to meet Microsoft, customer and/or government security screening requirements are required for this role. These requirements include but are not limited to the following specialized security screenings:
- Microsoft Cloud Background Check: This position will be required to pass the Microsoft Cloud Background Check upon hire/transfer and every two years thereafter.
Preferred Qualifications:
*
- B.S. Degree in Computer Engineering, Computer Science, Electrical or equivalent experience
- 10+ years of SSD firmware engineering development experience
- 6+ years of NVMe and PCIe experience
- Deep expertise in storage device resiliency, fault analysis, and live-site operations.
- Lead end-to-end design decisions across detection, prediction, mitigation, and repair of SSDs in hyper scale environment.
- Design component- reliability frameworks that work across different components
- Proven ability to build automation heavy systems that operate safely at hyperscale.
Hardware Engineering IC4 - The typical base pay range for this role across the U.S. is USD $119,800 - $234,700 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $160,200 - $261,000 per year.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role â technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
What Are The Top Skills Required For Azure Developers?
Highest Paying Tech Companies in Europe
Fullstack developer salary in Germany [2023]
7 Cloud Computing Trends Coming in 2025 for Developers