| Location | Seattle, WA |
| Industry | Retail |
| Company Size | 10,000 employees or more |
| Year Founded | 1994 |
| Website | http://Amazon.com/militaryroles |
AWS operates the world"s largest fleet of GPU-accelerated servers powering AI/ML training and inference at cloud scale. Our team defines the server architectures, drives the hardware designs, and owns the fleet quality for these platforms - from component selection through datacenter operations. If you want to shape the physical hardware that frontier models train on, this is the role.
We are seeking a Cloud Hardware Development Engineer to define server architectures based on workload demand, translate them into detailed component specifications, and drive validation from PCBA bring-up through rack integration. You will lead ODM design partners through development and production, triage hardware issues across manufacturing and datacenters, and own fleet quality metrics post-launch.
What You Will Do
You will define the hardware that runs the world"s largest AI training workloads. Your designs span thermal, mechanical, power, and signal integrity across GPU-accelerated platforms. You will drive validation from first silicon through fleet-scale deployment, triage failures correlating across PCIe, power delivery, memory, and accelerator interconnects, and feed root cause findings back into design improvements. When a new server platform launches at a large scale, the architecture, component choices, and quality gates are yours.
Why You Will Love It
The world"s most advanced frontier models train on the hardware you design. You will see your architecture decisions scale to a large fleet of servers. The team is deeply technical and high-trust - you own platforms end to end from architecture definition through fleet operations.
The Ideal Candidate
You think across the full hardware stack - from silicon packaging and power delivery to rack-level thermal and mechanical design. You are as comfortable reviewing a schematic as you are analyzing fleet failure data. You drive quality through data, not assumption, and you hold design partners to the same standard you hold yourself. You mentor and develop junior engineers, contribute to hiring, and share your expertise to make the team stronger.
Key job responsibilities
Architecture & Design
Validation & Bring-up
Fleet Quality & Continuous Improvement
Cross-Team Collaboration
May require occasional (<10%) regional and international travel to Design and Manufacturing Partner sites.
A day in the life
You start the day reviewing thermal and power validation data from an EVT build at your ODM partner. Mid-morning, you join a design review to close signal integrity findings on a high-speed accelerator interconnect. In the afternoon, you triage a fleet quality signal - correlating component-level failure data with manufacturing lot information to identify a systemic issue. You end the day aligning with architecture teams on requirements for the next-generation platform.
About the team
The Hardware Engineering AI/ML UltraServer platform team is a group of engineers and technical program managers directly responsible for launching GPU-accelerated servers into the AWS fleet. Located in Seattle, Austin, and Cupertino, we collaborate with global development teams and ODM partners to deliver next-generation AI/ML infrastructure deployed in datacenters worldwide. We move fast with small, empowered teams delivering end-to-end - from server conception through fleet-scale operations.