Amazon.com Inc logo

Systems Engineer, AWS Incident Response

Amazon.com Inc

  • Seattle, WA
  • 13 days ago

    Highlights

    Off-call, you work to reduce time-to-detection and time-to-mitigation: deep-diving recent events for detection that lagged the impact or alarms that trigger without customer impact, correcting the signal behind them, leading operational health reviews, driving other teams" corrective actions to completion, and owning projects that strengthen AIR"s detection and incident management, including automation you can build yourself. Between incidents you will obsess over metrics and detection analysis, building dashboards and mechanisms that surface problems before customers notice, and you will drive operational improvements that make the incident management ecosystem faster and more accurate.

    Numbers & Facts

    LocationSeattle, WA
    IndustryRetail
    Company Size10,000 employees or more
    Year Founded1994
    Websitehttp://Amazon.com/militaryroles

    Description

    As a Systems Engineer on AWS Incident Response, you will be on the front line of AWS incident response. You will lead high-severity calls, triage complex failures across distributed systems, coordinate resolver teams, and drive incidents to mitigation in real time while millions of customers depend on the outcome. Between incidents you will obsess over metrics and detection analysis, building dashboards and mechanisms that surface problems before customers notice, and you will drive operational improvements that make the incident management ecosystem faster and more accurate.

    You will make real-time decisions under pressure, deep-diving the largest and most complex technical environment in the world. You will develop expertise across AWS services, networking, and infrastructure, building a breadth of knowledge that few roles offer. Your scope spans all of AWS rather than a single service. You will own operational processes end to end and use data to find the next improvement in how we detect and mitigate faster. You will also have the opportunity to grow your development skills by taking on coding projects that accelerate incident response and reduce toil.

    This role includes participation in an on-call rotation covering weekdays, weekends, and holidays. On-call shifts fall within your local daytime hours.

    Key job responsibilities

    • Lead high-severity incident response calls end to end: assess impact, coordinate resolvers across AWS service teams, communicate clearly under pressure, manage escalations, and drive the incident to mitigation with documentation throughout.
    • Own and run operational health reviews, and build and maintain the dashboards, metrics, and monitoring that surface trends before they become incidents.
    • Improve detection accuracy and speed. Identify patterns across events and build proactive mechanisms that prevent recurrence.
    • Deep-dive operational data to find systemic issues, measure response effectiveness, and prioritize improvements against what the data shows is degrading.
    • Identify gaps in operational processes, documentation, and tooling, and build or improve mechanisms that reduce time-to-detection and time-to-mitigation.
    • Apply scripting, automation, and generative AI to accelerate incident response and reduce toil, including where AI can augment human judgment during an incident or surface insight from operational data at scale.
    • Work with service teams so that learnings from each incident drive corrective actions to completion, closing the loop between what broke and what gets fixed.
    • Mentor peers in your areas of technical and operational strength.

    A day in the life

    When you are on call, incidents take priority. You join the incident bridge, assess the scope of impact from real-time metrics and dashboards, engage the resolver teams that own the affected services, and drive the event to mitigation, escalating when progress stalls. Off-call, you work to reduce time-to-detection and time-to-mitigation: deep-diving recent events for detection that lagged the impact or alarms that trigger without customer impact, correcting the signal behind them, leading operational health reviews, driving other teams" corrective actions to completion, and owning projects that strengthen AIR"s detection and incident management, including automation you can build yourself.

    About the team

    AWS Incident Response (AIR), part of the AWS Resilience organization, ensures high availability of AWS. When major incidents hit, AIR leads the response, coordinating resolver teams across AWS and driving mitigation. We move fast, but not carelessly, obsessing over observability of the cloud and continuously improving our detection and response speed and accuracy. Each incident feeds improvements to our detection, processes, and tooling, making the next one shorter or preventing it entirely. This is a high-visibility, high-impact role with a global view of AWS health that few teams get to see.

    About Company

    At Amazon, we don’t wait for the next big idea to present itself. We envision the shape of impossible things and then we boldly make them reality. So far, this mindset has helped us achieve some incredible things. Let’s build new systems, challenge the status quo, and design the world we want to live in. We believe the work you do here will be the best work of your life.

    Wherever you are in your career exploration, Amazon likely has an opportunity for you. Our research scientists and engineers shape the future of natural language understanding with Alexa. Fulfillment center associates around the globe send customer orders from our warehouses to doorsteps. Product managers set feature requirements, strategy, and marketing messages for brand new customer experiences. And as we grow, we’ll add jobs that haven’t been invented yet.

    It’s Always Day 1
    At Amazon, it’s always “Day 1.” Now, what does this mean and why does it matter? It means that our approach remains the same as it was on Amazon’s very first day – to make smart, fast decisions, stay nimble, invent, and stay focused on delighting our customers. In our 2016 shareholder letter, Amazon CEO Jeff Bezos shared his thoughts on how to keep up a Day 1 company mindset. “Staying in Day 1 requires you to experiment patiently, accept failures, plant seeds, protect saplings, and double down when you see customer delight,” he wrote. “A customer-obsessed culture best creates the conditions where all of that can happen.” You can read the full letter here

    Our Leadership Principles
    Our Leadership Principles help us keep a Day 1 mentality. They aren’t just a pretty inspirational wall hanging. Amazonians use them, every day, whether they’re discussing ideas for new projects, deciding on the best solution for a customer’s problem, or interviewing candidates. To read through our Leadership Principles from Customer Obsession to Bias for Action, visit https://www.amazon.jobs/principles

    Similar Jobs

    See more jobs