Amazon, 2026
Agents that migrate AWS resources, and know when to stop
A multi-agent system for moving roughly a million SNS topics and 29 million SQS queues, with a human approval step before anything changes.
Amazon's Region Flex org needed to migrate a huge fleet of messaging resources: roughly a million SNS topics and 29 million SQS queues. If you don't live in AWS, think of a topic as a broadcast channel and a queue as a mailbox. Many of them belong to teams other than the one running the migration.
The hard part wasn't generating changes
At this scale, migrating by hand isn't realistic. But the obvious alternative, an agent that just goes and does it, has a trust problem. A wrong change to a resource another team depends on can break their service, and no one wants to hand that decision to a system they can't see into.
So I framed the problem as a governance question first and an automation question second: how do you get the speed of agents while keeping every change reviewable and every owner in control?
What I built
I built a multi-agent system with Kiro CLI, backed by a custom MCP server I wrote for it. The work runs through a seven-phase workflow that groups into three stages: generate, approve, execute.
- Stage 1
Generate
Agents inspect the resources and draft the migration as pull requests.
- Stage 2
Approve
The owning team reviews the change. Nothing moves without a yes.
- Stage 3
Execute
Only approved changes are applied.
I packaged the whole thing with AIM, Amazon's system for sharing AI agents, so any team can install it with one command and no new infrastructure. Adoption friction matters as much as capability when the users are busy engineers with their own roadmaps.
The moment it earned trust
In a real test, the system found a resource that was owned by a different team and refused to make the change. It became the highlight of the demo, because it showed the agents respected ownership instead of just following instructions.
Validating with real customers
Before scaling it up, I got on calls with teams who were actually migrating their SNS and SQS resources. I generated migration PRs for their resources with the agent, sent them over, and asked a simple question: is this what you'd want to see? Getting that sanity check early is much cheaper than finding out after a rollout.
This is the same habit I've had every summer at Amazon. Talking to users before building more is the cheapest way to find out you're wrong.
Outcome
The project won People's Top Choice at the Global AI Solutions Expo 2026.
The lesson I took from it is the one I'd bring to any AI product: people adopt automation when they can see what it's going to do and stop it if it's wrong. Capability gets the demo. Control gets the rollout.