> Markdown version of [/jobs/ext/3545414-cloud-hardware-dev-engineer-aws-ai-ml-ultraservers](https://www.wearedevelopers.com/jobs/ext/3545414-cloud-hardware-dev-engineer-aws-ai-ml-ultraservers). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Cloud Hardware Dev Engineer, AWS AI/ML UltraServers - **Company:** Amazon.com, Inc. - **Location:** Cupertino, CA, United States - **Experience:** Experienced - **Salary:** $157,300.0 - $212,800.0 - **Contract:** Permanent contract - **Skills:** Board Bringup, Artificial Intelligence, Amazon Web Services, Amazon Elastic Compute Cloud, Cloud Computing, Computer Engineering, Data Centers, Software Debugging, Firmware, Hardware Design, PCI Express, Signal Integrity, Strategies of Testing, Machine Learning Operations, Nvme - **Published:** October 1, 2026 - **Apply:** https://dejobs.org/x/x/CCED1CF2E6414F1ABAF25FDEC06E8A7D/job/ ## About the Role You think across the full hardware stack - from silicon packaging and power delivery to rack-level thermal and mechanical design. You are as comfortable reviewing a schematic as you are analyzing fleet failure data. You drive quality through data, not assumption, and you hold design partners to the same standard you hold yourself. You mentor and develop junior engineers, contribute to hiring, and share your expertise to make the team stronger., * Bachelor's degree in electrical engineering, computer engineering, or equivalent * Experience in server technologies such as, thermal, mechanical, power, and signal integrity * Experience in developing functional specifications, design verification plans and functional test procedures * 2+ years of hardware design, development and validation experience for server or compute platforms, * Master's degree in Electrical Engineering, Computer Engineering, or a related technical field * Experience with analog, digital, and high-speed circuit design * Experience working in data centers or critical infrastructure * 2+ years of experience working with ODMs through product development and manufacturing lifecycle * 2+ years experience working with hardware bring-up, debug, or validation of GPU/accelerator platforms * Familiarity with PCIe topology, NVMe, and accelerator interconnects * Experience developing and executing test procedures for electrical or mechanical systems ## Description Architecture & Design * Define server architectures based on workload demand and customer requirements, translating them into detailed designs and component specifications that enable high-performance AI training and inference at scale * Work with interdisciplinary teams of component, firmware, test, qualification, and integration engineers to deliver cohesive designs * Drive design reviews with ODM/JDM partners covering schematic, layout, BOM, and manufacturing DFx (Design for Test, Design for Manufacturing) Validation & Bring-up * Define and execute validation strategies from PCBA bring-up through server and rack integration - covering power sequencing, signal integrity, thermal characterization, and accelerator interconnect performance * Own hardware debug during EVT/DVT/PVT builds, correlating failures across PCIe, power rails, memory channels, and GPU subsystems * Triage hardware issues at both ODM facilities and datacenters, conduct root cause analysis, and implement corrective actions Fleet Quality & Continuous Improvement * Own fleet quality metrics post-launch: server-level annualized failure rates and component-level failure modes * Monitor operational telemetry to identify systemic issues and drive design or process changes for current and future platforms * Partner with test and automation teams to improve manufacturing yield and reduce test dwell times Cross-Team Collaboration * Work with EC2 architecture teams to align on instance definitions, workload requirements, and platform trade-offs * Drive ODM/JDM design partners through development milestones and production ramp * Collaborate with firmware, software, and operations teams to ensure designs are debuggable, serviceable, and automation-ready May require occasional (<10%) regional and international travel to Design and Manufacturing Partner sites. A day in the life You start the day reviewing thermal and power validation data from an EVT build at your ODM partner. Mid-morning, you join a design review to close signal integrity findings on a high-speed accelerator interconnect. In the afternoon, you triage a fleet quality signal - correlating component-level failure data with manufacturing lot information to identify a systemic issue. You end the day aligning with architecture teams on requirements for the next-generation platform. About the team The Hardware Engineering AI/ML UltraServer platform team is a group of engineers and technical program managers directly responsible for launching GPU-accelerated servers into the AWS fleet. Located in Seattle, Austin, and Cupertino, we collaborate with global development teams and ODM partners to deliver next-generation AI/ML infrastructure deployed in datacenters worldwide. We move fast with small, empowered teams delivering end-to-end - from server conception through fleet-scale operations. ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Playing Pong on a shoulder press machine](https://www.wearedevelopers.com/videos/100140-playing-pong-on-a-shoulder-press-machine) - [From Model to Metal: An Open Source Stack for Accelerating Intelligence](https://www.wearedevelopers.com/videos/1636-from-model-to-metal-an-open-source-stack-for-accelerating-intelligence) - [Agent Smith Gets Hardware: Autonomous IoT Hacking From Debug Port to Cloud API](https://www.wearedevelopers.com/videos/100258-agent-smith-gets-hardware-autonomous-iot-hacking-from-debug-port-to-cloud-api) - [Remote Driving on Plant Grounds with State-of-the-Art Cloud Technologies](https://www.wearedevelopers.com/videos/251-remote-driving-on-plant-grounds-with-state-of-the-art-cloud-technologies) - [My ongoing quest to code in VR](https://www.wearedevelopers.com/videos/2080-my-ongoing-quest-to-code-in-vr) ## Related Articles - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 162: AI careers, MCP, AWS best practices & floppy sweaters](https://www.wearedevelopers.com/magazine/571-dev-digest-162-ai-careers-mcp-aws-best-practices-floppy-sweaters) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Why Attend a Developer Event in 2026?](https://www.wearedevelopers.com/magazine/688-why-attend-a-developer-event-in-2026) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)