Cloud Hardware Dev Engineer (AWS Generative AI & ML Servers), AWS Hardware Engineering Services

Amazon.com, Inc.
Cupertino, CA, United States
3 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
2 years minimum
Compensation
$136,000.0 - $212,800.0
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Amazon Web Services Amazon Elastic Compute Cloud Amazon S3 Cloud Computing Databases Computer Engineering Data Centers Software Debugging File Server Firmware Hardware Design
+14 more
Machine Learning Productivity Software Systems Development Life Cycle Cloud Services Server Administration Software Engineering Subsystems AI Infrastructure Enterprise Software Applications High Performance Computing Large Language Models Generative AI Hardware Infrastructure Network Server

Job description

Do you want to build the backbone of Generative AI cloud at AWS? Do you want to build the future of the cloud for AI training and inference? Want to do industry leading work delivering continuous price performance improvements in the cloud for AI model training for multi billion variable LLMs? Come Join us in designing, delivering and operating AWS cloud offerings that enable high performance and scalability in AI/ML and HPC workloads.

Utility Computing (UC) AWS Utility Computing (UC) provides product innovations - from foundational services such as Amazon’s Simple Storage Service (S3) and Amazon Elastic Compute Cloud (EC2), to consistently released new product innovations that continue to set AWS’s services and features apart in the industry. As a member of the UC organization, you’ll support the development and management of Compute, Database, Storage, Internet of Things (IoT), Platform, and Productivity Apps services in AWS, including support for customers who require specialized security solutions for their cloud services.

You’ll join a diverse team of software, hardware, and network engineers, supply chain specialists, security experts, operations managers, and other vital roles. You’ll collaborate with people across AWS to help us deliver the highest standards for safety and security while providing seemingly infinite capacity at the lowest possible cost for our customers. And you’ll experience an inclusive culture that welcomes bold ideas and empowers you to own them to completion.

Key job responsibilities As a member of the Hardware Engineering Services team in this specific function, you will own and lead the design, development and root cause of a new segment of accelerated servers.

You will work closely with our customers to understand their technical needs and business goals, leveraging your experience with server design and the knowledge of various teams to architect the solutions that we will deploy at scale.

To deliver your products you will work with an interdisciplinary team of component, firmware, test, qualification, and integration engineers, and lead our design and manufacturing partners to bring these servers to the data center. After launch you will oversee the fleet of servers you develop, monitoring their quality and how they are meeting the customer requirements.

A day in the life Your day to day responsibilities will include interfacing with our internal and external customers to understand project requirements and facilitate system development ontop of your server design. You will be responsible for learning operational challenges to our existing fleet with the goal of improving the current customer experience as well as developing improved systems for future designs. You will work directly with vendors and ODM/JDM design teams to develop and manufacture your product at scale.

About the team The team is comprise of both Hardware Design Engineers, System Design Engineers, Software Development Engineers and Technical Program Managers, all with the common goal of delivering the best Accelerated Server fleet possible to our customers.

Requirements

  • Experience working with interdisciplinary teams to execute product design from concept to production
  • Experience developing and executing test procedures for mechanical or electrical systems/components based on design intent and approved equipment submissions
  • Knowledge of server hardware and components
  • Bachelor’s degree in electrical engineering or equivalent
  • 2+ years of server hardware troubleshooting and repair experience
  • 4+ years of hardware design and validation of components, subsystems and systems experience, * Master’s degree or above in electrical engineering, computer engineering, or equivalent
  • Experience in compute and storage server architecture and design for large scale applications
  • AI infrastructure hardware development and debugging experience

Benefits & conditions

Pulled from the full job description

  • AD&D insurance
  • Parental leave
  • Health insurance
  • 401(k) matching
  • Paid time off
  • Vision insurance
  • Dental insurance, The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits. USA, CA, Cupertino - 157,300.00 - 212,800.00 USD annually USA, WA, Seattle - 136,000.00 - 184,000.00 USD annually

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:19 min

Orchestrating over-the-air firmware updates for vehicle modules

Denis Grahovac · WWC 2021

3:43 min

The enduring legacy of the amazon S3 storage API

Chris Heilmann +3 · LIVE

3:04 min

Database evolution and the funding behind vector databases

Erik Bamberg · LIVE

1:12 min

Addressing the competitive landscape of specialized hardware demands

Hazal Mestci +1 · Coffee With Developers

4:01 min

Managing application isolation via pluggable database models

Wei Hu Wei Hu · WWC 2022

5:08 min

Automating data collection and managing crowdsourced training image sets

Kris Howard · LIVE

Videos

See all

Related articles

See all