Amazon.com Inc logo

System Development Engineer II, OIS Command Center

Amazon.com Inc

  • Nashville, TN
  • 6 days ago

    Highlights

    Most of your time goes to designing, building, and operating Reflex agents and the platform beneath them: the agent runtime and tool orchestration on Amazon Bedrock AgentCore, the LLM evaluation framework that gates each agent"s path from human-reviewed to autonomous, and the observability layer that keeps production agents accountable. You"ll design and deliver AI agents on Amazon Bedrock AgentCore that triage high-severity incidents, scribe live bridge calls in real time, draft stakeholder communications, and automate post-incident documentation and reporting, shifting incident management from a manual, pull-based model to an intelligent, push-based one.

    Numbers & Facts

    LocationNashville, TN
    IndustryRetail
    Company Size10,000 employees or more
    Year Founded1994
    Websitehttp://Amazon.com/militaryroles

    Description

    We"re seeking an experienced Principal Technical Program Manager to lead the Join us in building Reflex, the agentic incident management platform for Amazon"s fulfillment network. You"ll design and deliver AI agents on Amazon Bedrock AgentCore that triage high-severity incidents, scribe live bridge calls in real time, draft stakeholder communications, and automate post-incident documentation and reporting, shifting incident management from a manual, pull-based model to an intelligent, push-based one.

    The OIS Command Center (OCC) is Amazon"s 24/7 incident management function for high-severity incidents impacting fulfillment centers, delivery stations, and sortation centers worldwide, the infrastructure network that Amazon Robotics runs on. When this network degrades, robots stop and packages stop moving; OCC exists to make those minutes as short as possible. OCC manages roughly 1,500 high-severity incidents and triages some 14,000 alerts every year. Today, Incident Managers (IMs) continuously monitor signal feeds, engage resolver teams, and assemble a situational picture under time pressure before resolution work can even begin. Reflex changes that model fundamentally: agents watch the signals, assemble the context, and tell IMs when and how to engage, reserving human judgment for the decisions that actually need it.

    This is a builder role with an operational edge. Most of your time goes to designing, building, and operating Reflex agents and the platform beneath them: the agent runtime and tool orchestration on Amazon Bedrock AgentCore, the LLM evaluation framework that gates each agent"s path from human-reviewed to autonomous, and the observability layer that keeps production agents accountable. You"ll also periodically join live incident bridge calls in an Incident Manager capacity, staying close to the operational reality your software serves and turning what you learn on-call into what you build next. Your customers sit one Slack channel away, and you"ll experience the impact of what you ship on the very next incident call.

    Key job responsibilities

    • Design, build, test, and operate AI agents and supporting services on AWS (Amazon Bedrock AgentCore, serverless compute, event-driven pipelines) that automate incident triage, call scribing, communications, post-incident documentation, and operational reporting
    • Own features end-to-end: from sitting with Incident Managers to understand the workflow, through design, implementation, evaluation, deployment, and production operation
    • Build the platform foundations that gate agent autonomy, including LLM output evaluation, monitoring and alerting for agents in production, and identity and access controls aligned with Amazon standards
    • Design the feedback loops through which agents learn from Incident Managers: capturing reviews, corrections, and approvals as evaluation signal, and turning resolved incidents into structured history that improves pattern matching, severity classification, and resolver routing over time
    • Integrate Reflex with the incident ecosystem: ticketing, chat, telemetry, detection feeds, and live call transcription.
    • Raise the bar on operational excellence, security, and quality for AI systems acting inside production incident workflows

    A day in the life

    You might start by reviewing overnight agent evaluation results and tuning a tool integration before shipping an improvement IMs see on the next incident. Later, you pair with an Incident Manager to observe how they used the scribing agent on a live call, turning their corrections into evaluation signal that moves the agent closer to autonomous posting. You also build the platform, design feedback loops, and integrate with ticketing, chat, and detection feeds. And periodically, you take a seat on a high-severity bridge call as an Incident Manager, because the best way to know what to automate next is to carry the workload firsthand.

    Amazon offers a full range of benefits that support you and eligible family members, including domestic partners. Benefits can vary by location, the number of regularly scheduled hours you work, length of employment, and job status such as seasonal or temporary employment. The benefits that generally apply to regular, full-time employees include:

    1. Medical, Dental, and Vision Coverage

    2. Maternity and Parental Leave Options

    3. Paid Time Off (PTO)

    4. 401(k) Plan

    If you are not sure that every qualification on the list above describes you exactly, we"d still love to hear from you! At Amazon, we value people with unique backgrounds, experiences, and skillsets. If you're passionate about this role and want to make an impact on a global scale, please apply!

    About the team

    The OIS Command Center (OCC) is Amazon"s global, follow-the-sun incident management team within Operations Infrastructure Services, part of Amazon Robotics. OCC runs 24/7, managing roughly 1,500 high-severity incidents and triaging 14,000 alerts annually across fulfillment centers, delivery stations, and sortation centers worldwide. You"ll join a small, high-ownership engineering team whose charter is to transform OCC from manual monitoring and documentation toward agentic automation that lets Incident Managers focus on leading calls rather than manual correlation.

    About Company

    At Amazon, we don’t wait for the next big idea to present itself. We envision the shape of impossible things and then we boldly make them reality. So far, this mindset has helped us achieve some incredible things. Let’s build new systems, challenge the status quo, and design the world we want to live in. We believe the work you do here will be the best work of your life.

    Wherever you are in your career exploration, Amazon likely has an opportunity for you. Our research scientists and engineers shape the future of natural language understanding with Alexa. Fulfillment center associates around the globe send customer orders from our warehouses to doorsteps. Product managers set feature requirements, strategy, and marketing messages for brand new customer experiences. And as we grow, we’ll add jobs that haven’t been invented yet.

    It’s Always Day 1
    At Amazon, it’s always “Day 1.” Now, what does this mean and why does it matter? It means that our approach remains the same as it was on Amazon’s very first day – to make smart, fast decisions, stay nimble, invent, and stay focused on delighting our customers. In our 2016 shareholder letter, Amazon CEO Jeff Bezos shared his thoughts on how to keep up a Day 1 company mindset. “Staying in Day 1 requires you to experiment patiently, accept failures, plant seeds, protect saplings, and double down when you see customer delight,” he wrote. “A customer-obsessed culture best creates the conditions where all of that can happen.” You can read the full letter here

    Our Leadership Principles
    Our Leadership Principles help us keep a Day 1 mentality. They aren’t just a pretty inspirational wall hanging. Amazonians use them, every day, whether they’re discussing ideas for new projects, deciding on the best solution for a customer’s problem, or interviewing candidates. To read through our Leadership Principles from Customer Obsession to Bias for Action, visit https://www.amazon.jobs/principles

    Similar Jobs

    See more jobs