Amazon.com Inc logo

Sr Cloud Hardware Dev Engineer, AWS Generative AI & ML Servers

Amazon.com Inc

  • Seattle, WA
  • 3 days ago

    Highlights

    You will drive validation from first silicon through fleet-scale deployment, triage failures correlating across PCIe, power delivery, memory, and accelerator interconnects, and feed root cause findings back into design improvements. We are seeking a Cloud Hardware Development Engineer to define server architectures based on workload demand, translate them into detailed component specifications, and drive validation from PCBA bring-up through rack integration.

    Numbers & Facts

    LocationSeattle, WA
    IndustryRetail
    Company Size10,000 employees or more
    Year Founded1994
    Websitehttp://Amazon.com/militaryroles

    Description

    AWS operates the world"s largest fleet of GPU-accelerated servers powering AI/ML training and inference at cloud scale. Our team defines the server architectures, drives the hardware designs, and owns the fleet quality for these platforms - from component selection through datacenter operations. If you want to shape the physical hardware that frontier models train on, this is the role.

    We are seeking a Cloud Hardware Development Engineer to define server architectures based on workload demand, translate them into detailed component specifications, and drive validation from PCBA bring-up through rack integration. You will lead ODM design partners through development and production, triage hardware issues across manufacturing and datacenters, and own fleet quality metrics post-launch.

    What You Will Do

    You will define the hardware that runs the world"s largest AI training workloads. Your designs span thermal, mechanical, power, and signal integrity across GPU-accelerated platforms. You will drive validation from first silicon through fleet-scale deployment, triage failures correlating across PCIe, power delivery, memory, and accelerator interconnects, and feed root cause findings back into design improvements. When a new server platform launches at a large scale, the architecture, component choices, and quality gates are yours.

    Why You Will Love It

    The world"s most advanced frontier models train on the hardware you design. You will see your architecture decisions scale to a large fleet of servers. The team is deeply technical and high-trust - you own platforms end to end from architecture definition through fleet operations.

    The Ideal Candidate

    You think across the full hardware stack - from silicon packaging and power delivery to rack-level thermal and mechanical design. You are as comfortable reviewing a schematic as you are analyzing fleet failure data. You drive quality through data, not assumption, and you hold design partners to the same standard you hold yourself. You mentor and develop junior engineers, contribute to hiring, and share your expertise to make the team stronger.

    Key job responsibilities

    Architecture & Design

    • Define server architectures based on workload demand and customer requirements, translating them into detailed designs and component specifications that enable high-performance AI training and inference at scale
    • Work with interdisciplinary teams of component, firmware, test, qualification, and integration engineers to deliver cohesive designs
    • Drive design reviews with ODM/JDM partners covering schematic, layout, BOM, and manufacturing DFx (Design for Test, Design for Manufacturing)

    Validation & Bring-up

    • Define and execute validation strategies from PCBA bring-up through server and rack integration - covering power sequencing, signal integrity, thermal characterization, and accelerator interconnect performance
    • Own hardware debug during EVT/DVT/PVT builds, correlating failures across PCIe, power rails, memory channels, and GPU subsystems
    • Triage hardware issues at both ODM facilities and datacenters, conduct root cause analysis, and implement corrective actions

    Fleet Quality & Continuous Improvement

    • Own fleet quality metrics post-launch: server-level annualized failure rates and component-level failure modes
    • Monitor operational telemetry to identify systemic issues and drive design or process changes for current and future platforms
    • Partner with test and automation teams to improve manufacturing yield and reduce test dwell times

    Cross-Team Collaboration

    • Work with EC2 architecture teams to align on instance definitions, workload requirements, and platform trade-offs
    • Drive ODM/JDM design partners through development milestones and production ramp
    • Collaborate with firmware, software, and operations teams to ensure designs are debuggable, serviceable, and automation-ready

    May require occasional (<10%) regional and international travel to Design and Manufacturing Partner sites.

    A day in the life

    You start the day reviewing thermal and power validation data from an EVT build at your ODM partner. Mid-morning, you join a design review to close signal integrity findings on a high-speed accelerator interconnect. In the afternoon, you triage a fleet quality signal - correlating component-level failure data with manufacturing lot information to identify a systemic issue. You end the day aligning with architecture teams on requirements for the next-generation platform.

    About the team

    The Hardware Engineering AI/ML UltraServer platform team is a group of engineers and technical program managers directly responsible for launching GPU-accelerated servers into the AWS fleet. Located in Seattle, Austin, and Cupertino, we collaborate with global development teams and ODM partners to deliver next-generation AI/ML infrastructure deployed in datacenters worldwide. We move fast with small, empowered teams delivering end-to-end - from server conception through fleet-scale operations.

    About Company

    At Amazon, we don’t wait for the next big idea to present itself. We envision the shape of impossible things and then we boldly make them reality. So far, this mindset has helped us achieve some incredible things. Let’s build new systems, challenge the status quo, and design the world we want to live in. We believe the work you do here will be the best work of your life.

    Wherever you are in your career exploration, Amazon likely has an opportunity for you. Our research scientists and engineers shape the future of natural language understanding with Alexa. Fulfillment center associates around the globe send customer orders from our warehouses to doorsteps. Product managers set feature requirements, strategy, and marketing messages for brand new customer experiences. And as we grow, we’ll add jobs that haven’t been invented yet.

    It’s Always Day 1
    At Amazon, it’s always “Day 1.” Now, what does this mean and why does it matter? It means that our approach remains the same as it was on Amazon’s very first day – to make smart, fast decisions, stay nimble, invent, and stay focused on delighting our customers. In our 2016 shareholder letter, Amazon CEO Jeff Bezos shared his thoughts on how to keep up a Day 1 company mindset. “Staying in Day 1 requires you to experiment patiently, accept failures, plant seeds, protect saplings, and double down when you see customer delight,” he wrote. “A customer-obsessed culture best creates the conditions where all of that can happen.” You can read the full letter here

    Our Leadership Principles
    Our Leadership Principles help us keep a Day 1 mentality. They aren’t just a pretty inspirational wall hanging. Amazonians use them, every day, whether they’re discussing ideas for new projects, deciding on the best solution for a customer’s problem, or interviewing candidates. To read through our Leadership Principles from Customer Obsession to Bias for Action, visit https://www.amazon.jobs/principles

    Similar Jobs

    See more jobs