Software Development Manager (EC2 Nitro), EC2 Core Provisioning

Amazon.com Inc

Seattle, WA

JOB DETAILS
SKILLS
Amazon Elastic Compute Cloud (EC2), Automotive Repair and Maintenance, Best Practices, Computer Engineering, Computer Firmware, Continuous Deployment/Delivery, Continuous Integration, Debugging Skills, Distributed Computing, Hardware Components, Large-Scale Systems, Machine Learning, Mentoring, Multiplatform/Cross-Platform, People Management, Scripting (Scripting Languages), Software Development, Software Engineering, Team Lead/Manager, Team Player, Telemetry, Test Automation, Vehicle Fleets
LOCATION
Seattle, WA
POSTED
21 days ago

build, deploy, and manage applications with unparalleled flexibility and efficiency.

Join our dynamic team, where we apply agentic and machine-learning solutions to one of the hardest problems in the fleet: returning broken servers to production when there is no deterministic signal of what is wrong. You will build a learning, agent-driven decision engine on top of fleet telemetry and repair history, and you will ship it as a production software service that operates across millions of servers in every region, serving every EC2 line of business from core servers to accelerators and UltraServers. Sentinel is a direct lever on unsellable capacity and on the cost of running the fleet, and we are evolving it from human-authored decision rules into a system that recovers capacity on its own.

We are looking for an experienced Software Development Manager (SDM) to lead this team. The ideal candidate has led teams, thoroughly understands the design, development, and debugging of large-scale distributed systems, and is excited to apply ML and agentic techniques to real hardware-recovery problems. In this role, the manager will work with a broad group of technical teams across hardware, firmware, vetting, and provisioning.

Key job responsibilities

  • Lead and inspire a team of engineers, providing guidance, mentorship, and support to foster their professional growth.
  • Own the recovery decision engine that returns broken servers to sellable capacity, driving down unsellable rate and the time a host stays stuck. Take on the failures that have no deterministic signal, and evolve the engine from static, human-authored signatures into an agentic, ML-driven system that infers the right repair from fleet outcomes and improves with every recovery.
  • Build and operate this as a production software service - reliable, secure, and observable - running across millions of servers in every region, not a set of offline models or scripts.
  • Debug complex, system-level, multi-component failures across hardware, firmware, BMC, and the provisioning and vetting stack, and turn that diagnosis into automated, repeatable recovery.
  • Collaborate with hardware engineering, firmware, component owners, vetting, and provisioning teams to expand recovery coverage across platforms and drive failures upstream to their root cause so they stop recurring.
  • Raise the bar on the safety of autonomous action on production-bound capacity, holding a high security and operational standard for a service that runs across all regions, including restricted environments.
  • Champion best practices in software engineering, including code quality, testing, automation, and continuous integration and delivery (CI/CD).

About the Company

A

Amazon.com Inc

At Amazon, we don’t wait for the next big idea to present itself. We envision the shape of impossible things and then we boldly make them reality. So far, this mindset has helped us achieve some incredible things. Let’s build new systems, challenge the status quo, and design the world we want to live in. We believe the work you do here will be the best work of your life.

Wherever you are in your career exploration, Amazon likely has an opportunity for you. Our research scientists and engineers shape the future of natural language understanding with Alexa. Fulfillment center associates around the globe send customer orders from our warehouses to doorsteps. Product managers set feature requirements, strategy, and marketing messages for brand new customer experiences. And as we grow, we’ll add jobs that haven’t been invented yet.

It’s Always Day 1
At Amazon, it’s always “Day 1.” Now, what does this mean and why does it matter? It means that our approach remains the same as it was on Amazon’s very first day – to make smart, fast decisions, stay nimble, invent, and stay focused on delighting our customers. In our 2016 shareholder letter, Amazon CEO Jeff Bezos shared his thoughts on how to keep up a Day 1 company mindset. “Staying in Day 1 requires you to experiment patiently, accept failures, plant seeds, protect saplings, and double down when you see customer delight,” he wrote. “A customer-obsessed culture best creates the conditions where all of that can happen.” You can read the full letter here

Our Leadership Principles
Our Leadership Principles help us keep a Day 1 mentality. They aren’t just a pretty inspirational wall hanging. Amazonians use them, every day, whether they’re discussing ideas for new projects, deciding on the best solution for a customer’s problem, or interviewing candidates. To read through our Leadership Principles from Customer Obsession to Bias for Action, visit https://www.amazon.jobs/principles
COMPANY SIZE
10,000 employees or more
INDUSTRY
Retail
FOUNDED
1994
WEBSITE
http://Amazon.com/militaryroles