Senior UCIe PHY Expert, MLA Technology

Amazon.com, Inc.
Austin, TX, United States
about 1 month ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
3 years minimum
Compensation
$159,200.0 - $215,300.0
Working hours
Regular working hours
Job source

Tech stack

Adobe InDesign Systems Engineering Bash Shell Code Review Computer Engineering Data Centers Software Debugging Firmware Python (Programming Language) Lua (Scripting Language) Machine Learning Subsystems
+2 more
Information Technology Physical Design

Job description

As a senior member of the team, you will join a mixed group of hardware and software engineers working to design, integrate, and innovate the next generation of machine learning chips into Trainium servers. In this position it is expected that you will:

  • Collaborate with architects, design teams, and software engineers on our next generation ML chips

  • Support on-going debug and operations of previous ML chips within manufacturing and the data center

  • Dive deep into IP integration, packaging, silicon bring up, characterization, and validation of our UCIe subsystems

  • Independently develop the scripts you need to execute and collaborate with software engineers as your needs scale

A day in the life

A day in the life of a CHDE focused on UCIe on the MLA Technology team focuses on operational excellence, constructively identifying problems, prototyping solutions, and leading data collection at scale to improve our products. We start each day looking at our fleet, reviewing dashboards for emergent issues impacting our customers, partnering with other teams to drive complex debugs as it pertains to UCIe and associated SoC subsystems. We then look forward to the future technologies being developed and how we can best focus our efforts to help improve them and ensure a high quality product on behalf of our customers.

Our team members touch everything from electrical simulations, to hardware qualification on test benches, to software driven data center metrics, with a broad range of tasks across multiple skillsets where you can help improve the reliability and performance of our products. You help the team evolve by actively participating in design discussions, team planning, code reviews, tickets/metric reviews, and data center capacity initiatives. CHDEs on the MLA Technology Team are expected to help mentor others on the team in their area of expertise to help develop the team’s baseline skillsets and to participate in the hiring process for the team.

Requirements

Bachelor’s degree in Electrical Engineering, Computer Engineering, Systems Engineering, or related fields

  • 7+ years of experience in Silicon development

  • 3+ years in SOC/IO/Subsystems Experience working closely with physical design teams to develop highly optimized ASICs with excellent power, performance and area

  • Good understanding of UCIe at the PHY and controller level

  • Good knowledge of UCIe training, timing parameters and/or controller features

  • Support the physical design team with IP integration, 2.5D packaging, clocking and timing constraints

  • Ability to create scripts (lua, bash, python, etc.) to accomplish functional day to day tasks.

  • Drive cross-functional triage effort on functional and performance issues

  • Take the leadership role in post-silicon bring-up of UCIe-Advanced or Standard

  • Perform system-level debug and root-cause analysis through bring-up, characterization, validation and production phases

Preferred Qualifications

  • MS degree in computer science, electrical engineering, or related field

  • Strong Firmware development skills within embedded environments

  • Good leadership skills and ability to multi-task and thrive in a dynamic environment

  • Knowledge of UCIe phy and controller related protocols

  • Good communication skills and interpersonal skills

Benefits & conditions

The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits .

USA, TX, Austin - 159,200.00 - 215,300.00 USD annually

About the company

Annapurna Labs (our organization within AWS UC) designs silicon and software that accelerates innovation. Customers choose us to create cloud solutions that solve challenges that were unimaginable a short time ago-even yesterday. Our custom chips, accelerators, and software stacks enable us to take on technical challenges that have never been seen before, and deliver results that help our customers change the world.

In Annapurna Labs we are at the forefront of hardware/software co-design not just in Amazon Web Services (AWS) but across the industry. Our Machine Learning Accelerator (MLA) Technology is seeking a UCIe PHY expert who is interested in diving deep into the definition, design, validation, and data center operation of AWS’s next generation machine learning silicon and servers.

As a senior member of our technology team, you will have opportunities to participate in the design and execution of UCIe, SERDES, and general high speed analog technologies, with the goal of creating the most stable machine learning platforms within AWS’s data centers. A senior UCIE engineer on our team needs to be able to work with vendors and internal design teams, understand UCIe timings and features, write/modify tests at scale, debug fleet wide issues, and collect data from manufacturing and the data center.

Our broader team has end to end ownership of some of the most complicated IPs on the most advanced server hardware in the world. We drive complex technical debug efforts involving our IPs and leverage the massive scale of EC2 to monitor, optimize, and improve our machine learning hardware reliability on behalf of our customers.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.juju.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:19 min

Orchestrating over-the-air firmware updates for vehicle modules

Denis Grahovac · WWC 2021

3:52 min

Avoiding remote code execution from unsanitized inputs

Alexander Pirker · WWC 2022

2:05 min

Finding product boundaries using lack of cohesion metric grouping

Pratishtha Pandey Pratishtha Pandey · WWC 2024

4:18 min

Prioritizing communication and structural awareness over strict tool mastery

Liam Hurrel +1 · WWC 2021

2:20 min

Utilizing custom firmware for variable torque manipulation

Daniel Meilak Daniel Meilak +1 · WWC Europe 2026

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

Videos

See all

Related articles

See all