Amazon.com Inc logo

Worldwide Specialist Solutions Architect - AI, Data & AI GTM

Amazon.com Inc

  • Chicago, IL
  • 1 day ago

    Highlights

    Your day starts with a whiteboard session alongside a financial services customer"s ML engineering team, walking them through a distributed fine-tuning architecture - helping them set up Supervised Fine-Tuning (SFT) with LoRA on a Llama model using SageMaker Training Jobs across a cluster of P5e instances, configuring FSDP for efficient multi-GPU parallelism, and advising them on checkpointing strategies so they can resume training without losing hours of compute. Solution Design & Delivery: Partner with customers" data science and engineering teams to deeply understand their business objectives, then architect solutions that leverage AWS AI/ML services - with emphasis on model customization (fine-tuning, continued pre-training, distillation) and inference optimization (model compilation, quantization, endpoint auto-scaling, multi-model endpoints).

    Numbers & Facts

    LocationChicago, IL
    IndustryRetail
    Company Size10,000 employees or more
    Year Founded1994
    Websitehttp://Amazon.com/militaryroles

    Description

    Generative AI and large-scale machine learning are redefining what"s possible - and AWS is at the center of that transformation. We are looking for a Machine Learning Solutions Architect (ML SA) who will serve as the technical authority on model customization and inference to help customers across the AMERICAS unlock the full potential of foundation models, custom training, and production-scale serving on AWS.

    Amazon has invested in AI for over two decades. From the recommendation engines that power Amazon.com to the deep learning behind Alexa, Prime Air, Amazon Go, and our supply chain optimization - machine learning is embedded in everything we build. Now, through Amazon SageMaker AI, SageMaker HyperPod, Amazon Bedrock, and our purpose-built silicon (Trainium, Inferentia), we are enabling customers to fine-tune, train, and deploy models at unprecedented scale and efficiency.

    As a Model Customization & Inference SageMaker ML SA, you will work directly with customers - from startups to enterprises - to design end-to-end ML architectures that span the full lifecycle: data preparation, distributed training, model fine-tuning (LoRA, PEFT, RLHF), inference optimization, and production deployment. You will operate across all 2 layers of the AWS AI/ML stack:

    Infrastructure & Compute - SageMaker HyperPod, GPU-based EC2, EKS/ECS for ML and Gen AI workloads

    ML Platforms - Amazon SageMaker AI (training jobs, endpoints, pipelines, MLOps)

    You will be the bridge between customers and AWS engineering - translating real-world business problems into scalable ML architectures and feeding critical customer signals back to service teams to shape the product roadmap.

    Key job responsibilities

    Solution Design & Delivery: Partner with customers" data science and engineering teams to deeply understand their business objectives, then architect solutions that leverage AWS AI/ML services - with emphasis on model customization (fine-tuning, continued pre-training, distillation) and inference optimization (model compilation, quantization, endpoint auto-scaling, multi-model endpoints).

    Technical Leadership: Serve as the go-to SME on model customization and inference patterns across SageMaker AI and SageMaker HyperPod. Guide field SAs and customers on best practices for training at scale and deploying models with optimal latency, throughput, and cost.

    Customer Adoption & Revenue Impact: Partner with Specialist SAs, Account Teams, Sales, and Business Development to accelerate adoption of SageMaker AI across the AMERICAS - directly contributing to pipeline generation, opportunity progression, and revenue attainment.

    Thought Leadership & Evangelism: Author technical blogs, whitepapers, reference architectures, and reusable solution artifacts. Deliver presentations at flagship events (AWS re:Invent, AWS Summits, industry conferences) to establish AWS as the leader in model customization and inference.

    Voice of the Customer: Act as the technical liaison between customers and AWS service teams (SageMaker). Capture and escalate product feature requests, identify gaps, and drive platform improvements grounded in real-world customer needs.

    Community Building: Develop and scale an internal community of ML subject matter experts across the AMERICAS, fostering knowledge sharing on model customization, inference optimization, and emerging ML patterns.

    A day in the life

    Your day starts with a whiteboard session alongside a financial services customer"s ML engineering team, walking them through a distributed fine-tuning architecture - helping them set up Supervised Fine-Tuning (SFT) with LoRA on a Llama model using SageMaker Training Jobs across a cluster of P5e instances, configuring FSDP for efficient multi-GPU parallelism, and advising them on checkpointing strategies so they can resume training without losing hours of compute. By midday, you"re on a call with a retail customer who"s struggling with inference latency on their real-time recommendation model - you dig into their endpoint configuration, recommend migrating to a SageMaker real-time inference endpoint backed by GPU or custom chips,and help them benchmark quantized vs. full-precision serving to hit their P99 latency targets. After lunch, you carve out time to author a reference architecture blog on multi-model endpoints for generative AI workloads, review a Product Feature Request (PFR) you"re drafting based on customer feedback around SageMaker HyperPod training job scheduling, and jump into a Slack thread with the service team to advocate for a customer-requested enhancement to Serverless Model Customization. You close the day prepping a re:Invent chalk talk on cost-optimized inference patterns - pulling real customer benchmarks, tuning your demo notebook, and syncing with your coverage SA on an upcoming Bedrock-to-SageMaker migration opportunity that could unlock a six-figure SageMaker pipeline. No two days look the same, but every day centers on one thing: helping customers get from raw model to production-grade, cost-efficient inference - faster.

    About the team

    The Worldwide Specialist Organization (WWSO) SageMaker AI team is a group of deeply technical Solutions Architects, Data Scientists, and ML Engineers who serve as the global technical authority on Amazon SageMaker AI. We sit at the intersection of customers and product - working hands-on with enterprises across every industry to design and deliver end-to-end ML solutions spanning model customization, distributed training, inference optimization, and MLOps at scale. Our charter is threefold: build reusable reference architectures and solutions that act as force multipliers for the field, drive specialist customer engagements on the most complex and high-impact ML workloads, and shape the SageMaker AI product roadmap by translating real-world customer signals into product priorities.

    About Company

    At Amazon, we don’t wait for the next big idea to present itself. We envision the shape of impossible things and then we boldly make them reality. So far, this mindset has helped us achieve some incredible things. Let’s build new systems, challenge the status quo, and design the world we want to live in. We believe the work you do here will be the best work of your life.

    Wherever you are in your career exploration, Amazon likely has an opportunity for you. Our research scientists and engineers shape the future of natural language understanding with Alexa. Fulfillment center associates around the globe send customer orders from our warehouses to doorsteps. Product managers set feature requirements, strategy, and marketing messages for brand new customer experiences. And as we grow, we’ll add jobs that haven’t been invented yet.

    It’s Always Day 1
    At Amazon, it’s always “Day 1.” Now, what does this mean and why does it matter? It means that our approach remains the same as it was on Amazon’s very first day – to make smart, fast decisions, stay nimble, invent, and stay focused on delighting our customers. In our 2016 shareholder letter, Amazon CEO Jeff Bezos shared his thoughts on how to keep up a Day 1 company mindset. “Staying in Day 1 requires you to experiment patiently, accept failures, plant seeds, protect saplings, and double down when you see customer delight,” he wrote. “A customer-obsessed culture best creates the conditions where all of that can happen.” You can read the full letter here

    Our Leadership Principles
    Our Leadership Principles help us keep a Day 1 mentality. They aren’t just a pretty inspirational wall hanging. Amazonians use them, every day, whether they’re discussing ideas for new projects, deciding on the best solution for a customer’s problem, or interviewing candidates. To read through our Leadership Principles from Customer Obsession to Bias for Action, visit https://www.amazon.jobs/principles

    Similar Jobs