Amazon.com Inc logo

Sr Technical Program Manager - Hardware, AWS Generative AI & ML Servers

Amazon.com Inc

  • Seattle, WA
  • 2 days ago

    Highlights

    You will translate ambiguity into structure: turning a fleet telemetry signal into a corrective action plan with quantified failure rates, a customer requirement into a new platform milestone with EVT/DVT/PVT gates, or a manufacturing escape into a design change with updated validation criteria. Build program timelines aligned to different phases (Program Initiation, Design, Qualification, Pilot, Post-launch) with critical path analysis, risk identification with new & unique changes to the hardware, and milestone tracking for the program.

    Numbers & Facts

    LocationSeattle, WA
    IndustryRetail
    Company Size10,000 employees or more
    Year Founded1994
    Websitehttp://Amazon.com/militaryroles

    Description

    AWS operates the world"s largest fleet of GPU-accelerated servers powering AI/ML workloads at cloud scale. Our team designs, builds, and operates this fleet - solving systemic hardware issues and building systems that detect and prevent recurrence so customers experience the highest quality of service.

    We are seeking a Senior Technical Program Manager to drive end-to-end delivery of GPU-accelerated servers across our global fleet. You will coordinate cross-functional engineering teams spanning hardware, firmware, and software, manage ODM partnerships across multiple continents, and establish closed-loop quality systems that drive continuous improvements. This role requires technical depth to translate engineering constraints into program risk, combined with program management excellence to deliver complex hardware at global scale.

    What You Will Do

    You will own programs where the critical path runs through silicon, firmware, and software teams simultaneously. You will translate ambiguity into structure: turning a fleet telemetry signal into a corrective action plan with quantified failure rates, a customer requirement into a new platform milestone with EVT/DVT/PVT gates, or a manufacturing escape into a design change with updated validation criteria. You will drive decisions on program trade-offs - adjusting scope when qualification gates slip, balancing deployment speed against fleet risk, and determining when to accept interim mitigations instead of holding for root-cause fixes. When a large scale of GPU servers depend on your program landing on time, you are the one ensuring hardware readiness, qualification completeness, and operational handoff happen without gaps.

    Why You Will Love It

    The world"s most advanced frontier models are trained on the platforms you help build. Your programs launch the GPU servers that power the largest AI/ML workloads on the planet. You will see your decisions reflected in fleet reliability metrics within weeks of deployment. The team is small and high-trust - you own programs end to end from concept through production, with direct access to leadership and engineering alike.

    The Ideal Candidate

    You have deep technical intuition across hardware and software - enough to challenge engineering decisions, not just track them. You thrive in ambiguity, bringing structure to programs where requirements, timelines, and dependencies are still forming. You align priorities across teams in different organizations, and you escalate with data, not noise. You actively mentor and develop others - TPMs and engineers alike. You contribute to hiring, promotion assessments, and raising the bar for program management practices in your organization.

    Key job responsibilities

    Strategy & Mechanisms

    • Define program strategy, objectives, and success criteria; influence resource allocation and priority decisions across engineering workstreams to align with organizational goals
    • Build and own mechanisms for program visibility - defining metrics, dashboards, and review cadences that enable data-driven decisions and early risk detection
    • Streamline delivery processes across teams; identify and eliminate dependencies, redundant gates, or coordination overhead that slow velocity

    Requirements & Planning

    • Facilitate requirements gathering with internal customers; develop Technical Requirements Documents (TRDs) covering server specs, rack configurations, PCIe topology, power/cooling topology, and SKU definitions
    • Build program timelines aligned to different phases (Program Initiation, Design, Qualification, Pilot, Post-launch) with critical path analysis, risk identification with new & unique changes to the hardware, and milestone tracking for the program.
    • Drive Program Initiation reviews: scope definition, preliminary annualized failure rate predictions, resource planning, supply chain long-lead identification, and RFP (Request for Proposal) issuance to ODMs

    Execution & Coordination

    • Drive cross-functional alignment across hardware, firmware (BIOS, BMC, CPLD), software, and operations teams through design reviews, manufacturing readiness and production readiness for fleet deployment.
    • Manage ODM partnerships: track EVT/DVT/PVT builds, manufacturing readiness gates, Bill of Materials (BOM) management in PLM systems, and quality checkpoints
    • Identify blockers early, escalate dependencies before they impact critical path, and facilitate technical trade-off decisions across engineering workstreams.
    • Communicate program status to leadership with clear reporting on milestone progress, risk posture, and mitigation plans.

    Risk & Quality

    • Challenge technical workstreams to surface risks early; quantify impact to schedule and reliability (annualized failure rate targets, availability SLAs) with proposed mitigations
    • Drive root cause analysis of fleet-wide hardware failures and ensure corrective actions flow back into qualification criteria and design requirements
    • Define acceptance criteria, coordinate qualification testing at server and rack levels, and manage go/no-go decisions for mission-critical AI/ML workloads

    Transition

    • Conduct knowledge transfer and document lessons learned; ensure operational readiness including automation, monitoring, and runbook completeness for production handoff

    May require occasional (<10%) regional and international travel to Design and Manufacturing Partner sites.

    A day in the life

    You start the day syncing with ODM partners across time zones on build status and open engineering actions. Mid-morning, you run an engineering review connecting firmware, software, and hardware teams to unblock a qualification gate. In the afternoon, you triage a fleet reliability issue with operations data, drive alignment on corrective actions, and update executive stakeholders on program risk posture. You end the day reviewing NPI milestone readiness and ensuring the next design review has clear entry criteria.

    About the team

    The Hardware Engineering AI/ML UltraServer platform team is a group of engineers and technical program managers directly responsible for launching GPU-accelerated servers into the AWS fleet. Located in Seattle, Austin, and Cupertino, we collaborate with global development teams and ODM partners to deliver next-generation AI/ML infrastructure deployed in datacenters worldwide. We move fast with small, empowered teams delivering end-to-end - from server conception through fleet-scale operations.

    About Company

    At Amazon, we don’t wait for the next big idea to present itself. We envision the shape of impossible things and then we boldly make them reality. So far, this mindset has helped us achieve some incredible things. Let’s build new systems, challenge the status quo, and design the world we want to live in. We believe the work you do here will be the best work of your life.

    Wherever you are in your career exploration, Amazon likely has an opportunity for you. Our research scientists and engineers shape the future of natural language understanding with Alexa. Fulfillment center associates around the globe send customer orders from our warehouses to doorsteps. Product managers set feature requirements, strategy, and marketing messages for brand new customer experiences. And as we grow, we’ll add jobs that haven’t been invented yet.

    It’s Always Day 1
    At Amazon, it’s always “Day 1.” Now, what does this mean and why does it matter? It means that our approach remains the same as it was on Amazon’s very first day – to make smart, fast decisions, stay nimble, invent, and stay focused on delighting our customers. In our 2016 shareholder letter, Amazon CEO Jeff Bezos shared his thoughts on how to keep up a Day 1 company mindset. “Staying in Day 1 requires you to experiment patiently, accept failures, plant seeds, protect saplings, and double down when you see customer delight,” he wrote. “A customer-obsessed culture best creates the conditions where all of that can happen.” You can read the full letter here

    Our Leadership Principles
    Our Leadership Principles help us keep a Day 1 mentality. They aren’t just a pretty inspirational wall hanging. Amazonians use them, every day, whether they’re discussing ideas for new projects, deciding on the best solution for a customer’s problem, or interviewing candidates. To read through our Leadership Principles from Customer Obsession to Bias for Action, visit https://www.amazon.jobs/principles

    Similar Jobs

    See more jobs