Amazon.com Inc logo

Senior Performance Engineer, Efficiency Red Team

Amazon.com Inc

  • Seattle, WA
  • 8 days ago

    Highlights

    You will work at the intersection of hardware and software, from GPU memory management in generative AI inference pipelines to DRAM allocation patterns deep inside language runtime, from storage subsystem inefficiencies to CPU scheduling behaviors at hyper-scale. Your day could include profiling a widely deployed library to quantify its memory allocation overhead, then building a prototype that demonstrates a significant reduction in DRAM footprint and calculating the impact in freed servers and dollars.

    Numbers & Facts

    LocationSeattle, WA
    IndustryRetail
    Company Size10,000 employees or more
    Year Founded1994
    Websitehttp://Amazon.com/militaryroles

    Description

    Are you the kind of engineer who obsesses over performance? Does the idea of hunting for hidden waste across the world"s largest infrastructure and turning those findings into hundreds of millions of dollars in freed capacity sound like the most exciting job you can imagine? If so, keep reading.

    This role offers something rare. Deep technical research freedom, immediate access to the largest compute infrastructure in the world, and direct line of sight to impact that reshapes how Amazon builds and operates its infrastructure. The Efficiency Red Team has already driven billions in cumulative capacity savings, reclaimed petabytes of DRAM through automated profiling, and built GPU efficiency frameworks adopted across Amazon. And we are just getting started. The surface area of opportunity grows with every new workload, every new chip generation, and every new AI model Amazon deploys. You would be joining a team with a proven foundation and an expanding frontier.

    Compute demand is rising with no sign of slowing down. The explosive growth of Generative AI has created unprecedented demand for GPUs, and the shock-waves are now constraining DRAM availability for traditional compute across the entire industry. Silicon, power, and physical space are not infinite. When capacity is constrained, efficiency becomes the new capacity. At Amazon"s scale, a single efficiency pattern discovered and applied can unlock resources equivalent to building entirely new data centers. A modest reduction in memory footprint across a widely deployed library translates into millions of dollars in freed capacity. An optimization that improves response latency by milliseconds across billions of requests directly improves the experience of hundreds of millions of customers and drives business growth.

    We are looking for experienced performance engineers to join a focused group whose mission is discovering and eliminating waste across every layer of the compute stack. You will work at the intersection of hardware and software, from GPU memory management in generative AI inference pipelines to DRAM allocation patterns deep inside language runtime, from storage subsystem inefficiencies to CPU scheduling behaviors at hyper-scale.

    Key job responsibilities

    • Apply first-principles analysis across CPU, DRAM, GPU, storage, I/O, and networking layers to uncover optimization opportunities that others overlook
    • Design and lead rigorous investigations that establish root causes and produce actionable findings with broad impact
    • Build proof-of-concept implementations for novel efficiency approaches such as memory compression, smart over-subscription, and language-level rewrites
    • Develop repeatable patterns and practices that translate individual findings into fleet-wide optimization playbooks
    • Quantify business impact at Amazon scale, connecting every technical insight to capacity freed, dollars saved, and customer experience improved
    • Collaborate across organizational boundaries where the largest opportunities often live at the intersection of teams, systems, and technology stacks
    • Contribute findings to internal profiling tools, automated optimization agents, and engineering guidance that become force multipliers across Amazon
    • Directly influence how Amazon reasons about performance engineering and capacity efficiency at scale as a founding member of the Efficiency Red Team

    A day in the life

    Your day could include profiling a widely deployed library to quantify its memory allocation overhead, then building a prototype that demonstrates a significant reduction in DRAM footprint and calculating the impact in freed servers and dollars. You might be deep in GPU utilization data, investigating why a generative AI training cluster shows substantial idle time during peak hours and designing an over-subscription model that safely reclaims that capacity. You could be translating complex performance findings into a clear narrative for senior leadership, showing them that a single optimization pattern applied across Amazon"s fleet frees capacity worth more than most companies spend on infrastructure in a year.

    Some weeks you will be reading CPU performance counters and cache miss rates. Other weeks you will be analyzing token throughput economics in large language model serving infrastructure. The variety is deliberate because waste hides everywhere, and finding it requires curiosity that refuses to stay in a single lane.

    About the team

    The Efficiency Red Team is a small group of performance engineers who move fluidly across the entire compute stack. The same engineer might profile CPU cache behavior in a traditional workload one week and analyze token throughput in a GPU inference cluster the next. You will have the freedom to pursue deep technical research, choose your own investigations, and follow the data wherever it leads. This is a founding opportunity to help shape the mission, methods, and culture of a team designed to be Amazon"s center of excellence in performance engineering. Every pattern this team discovers feeds directly into detection and remediation tools that operate continuously across the fleet, turning individual insights into lasting, compounding impact.

    About Company

    At Amazon, we don’t wait for the next big idea to present itself. We envision the shape of impossible things and then we boldly make them reality. So far, this mindset has helped us achieve some incredible things. Let’s build new systems, challenge the status quo, and design the world we want to live in. We believe the work you do here will be the best work of your life.

    Wherever you are in your career exploration, Amazon likely has an opportunity for you. Our research scientists and engineers shape the future of natural language understanding with Alexa. Fulfillment center associates around the globe send customer orders from our warehouses to doorsteps. Product managers set feature requirements, strategy, and marketing messages for brand new customer experiences. And as we grow, we’ll add jobs that haven’t been invented yet.

    It’s Always Day 1
    At Amazon, it’s always “Day 1.” Now, what does this mean and why does it matter? It means that our approach remains the same as it was on Amazon’s very first day – to make smart, fast decisions, stay nimble, invent, and stay focused on delighting our customers. In our 2016 shareholder letter, Amazon CEO Jeff Bezos shared his thoughts on how to keep up a Day 1 company mindset. “Staying in Day 1 requires you to experiment patiently, accept failures, plant seeds, protect saplings, and double down when you see customer delight,” he wrote. “A customer-obsessed culture best creates the conditions where all of that can happen.” You can read the full letter here

    Our Leadership Principles
    Our Leadership Principles help us keep a Day 1 mentality. They aren’t just a pretty inspirational wall hanging. Amazonians use them, every day, whether they’re discussing ideas for new projects, deciding on the best solution for a customer’s problem, or interviewing candidates. To read through our Leadership Principles from Customer Obsession to Bias for Action, visit https://www.amazon.jobs/principles

    Similar Jobs

    See more jobs