Amazon.com Inc logo

Sr. Software Development Engineer - AI/ML Networking Disaggregated Inference, Annapurna Labs, Annapurna Labs

Amazon.com Inc

  • Cupertino, CA
  • 9 days ago

    Highlights

    We"re looking for an engineer to work at the frontier of disaggregated inference: splitting LLM serving into separate prefill and decode pools and moving the model"s KV cache between them at the absolute limit of what the hardware allows. Experience with high-speed networking or HPC interconnects (RDMA, InfiniBand, libfabric, UCX, NIXL, MPI) is valued highly; embedded-systems experience is a plus.

    Numbers & Facts

    LocationCupertino, CA
    IndustryRetail
    Company Size10,000 employees or more
    Year Founded1994
    Websitehttp://Amazon.com/militaryroles

    Description

    Every token a large language model generates depends on data reaching the right accelerator at the right moment. As AI models outgrow any single chip, the network between accelerators becomes the bottleneck that decides how fast - and how affordably - the world"s largest models can serve real users. That network layer is what our team builds.

    We"re looking for an engineer to work at the frontier of disaggregated inference: splitting LLM serving into separate prefill and decode pools and moving the model"s KV cache between them at the absolute limit of what the hardware allows. Get it right and users get answers in milliseconds; get it wrong and the fastest accelerators in the world sit idle waiting on data. You"ll own pieces of the high-speed transfer path that make the difference, and you"ll measure your success in how close you run to the theoretical peak of the machine.

    In this role you will:

    Build and optimize the low-level data-movement software across accelerators, servers, and heterogeneous memory - over AWS"s highest-performance network fabric.

    Push performance to the hardware roofline: profile, find the real bottleneck, and close the gap between "it works" and "it runs as fast as physics permits."

    Work across the stack - from kernel and network transport up to the inference frameworks - and partner with teams building the chips, runtime, and models.

    Deliver features that run on our largest clusters, for our largest customers, serving the largest AI models in production.

    What we"re looking for:

    Strong C/C++ and a love for low-level, performance-critical systems - solid command of Linux, kernels, memory, and writing fast code.

    The instinct to ask "how fast could this possibly go?" and the rigor to measure it.

    Experience with high-speed networking or HPC interconnects (RDMA, InfiniBand, libfabric, UCX, NIXL, MPI) is valued highly; embedded-systems experience is a plus.

    Prior AI/ML experience is welcome but not required - if you"re a great systems engineer, we"ll teach you the ML side.

    If you like solving genuinely hard problems, working shoulder-to-shoulder with HPC and ML customers, iterating fast, and shipping at a scale few places can offer, come join us. This is a role on the leading edge of AI/ML infrastructure.

    About the team: You"d be joining Annapurna Labs, an integral part of AWS. Annapurna designs the hardware and software building blocks behind EC2 - every EC2 instance runs on hardware we designed. We specialize in the chips, systems, and software that optimize the AWS customer experience, and this team sits where cutting-edge AI meets the silicon and the network underneath it.

    A day in the life

    Annapurna Labs, a crucial part of AWS, is responsible for developing hardware and software components for EC2 infrastructure. Our team focuses on building networking solutions that for Machine Learning (ML) and High-Performance Computing (HPC) workloads on AWS.

    We have mixed discipline orgs, you'd be working side by side with infrastructure experts, hardware engineers, RTL engineers, scientists & architects. Our workforce spans the globe and is truly international, you'll find yourself working side by side with individuals from numerous countries. We take mentorship seriously, you can both expect senior mentorship and will be expected to mentor new and junior engineers.

    The pace is fast as we work on the latest advancements of AI/ML, but we take the time to bond as a team and enjoy the successes. We offer flexibility in working hours, and respect WLB as a core org tenet. The team enjoys working with numerous principal-level engineers and closely with directors, career growth opportunities are certainly available. This is a role where you will always be encouraged to keep learning, the AI/ML field is fast moving and constantly evolving.

    About Company

    At Amazon, we don’t wait for the next big idea to present itself. We envision the shape of impossible things and then we boldly make them reality. So far, this mindset has helped us achieve some incredible things. Let’s build new systems, challenge the status quo, and design the world we want to live in. We believe the work you do here will be the best work of your life.

    Wherever you are in your career exploration, Amazon likely has an opportunity for you. Our research scientists and engineers shape the future of natural language understanding with Alexa. Fulfillment center associates around the globe send customer orders from our warehouses to doorsteps. Product managers set feature requirements, strategy, and marketing messages for brand new customer experiences. And as we grow, we’ll add jobs that haven’t been invented yet.

    It’s Always Day 1
    At Amazon, it’s always “Day 1.” Now, what does this mean and why does it matter? It means that our approach remains the same as it was on Amazon’s very first day – to make smart, fast decisions, stay nimble, invent, and stay focused on delighting our customers. In our 2016 shareholder letter, Amazon CEO Jeff Bezos shared his thoughts on how to keep up a Day 1 company mindset. “Staying in Day 1 requires you to experiment patiently, accept failures, plant seeds, protect saplings, and double down when you see customer delight,” he wrote. “A customer-obsessed culture best creates the conditions where all of that can happen.” You can read the full letter here

    Our Leadership Principles
    Our Leadership Principles help us keep a Day 1 mentality. They aren’t just a pretty inspirational wall hanging. Amazonians use them, every day, whether they’re discussing ideas for new projects, deciding on the best solution for a customer’s problem, or interviewing candidates. To read through our Leadership Principles from Customer Obsession to Bias for Action, visit https://www.amazon.jobs/principles

    Similar Jobs

    See more jobs