Back to posts
AI & Productivity•9 min read

Harry Potter and the Order of the Agents

Cris Ryan Tan

Cris Ryan Tan

Senior Software Engineer

I already used several AI agents, each for a different kind of work, but they didn't work as a fleet. Each one ran its own long session, nothing managed how many tokens they burned, and nothing automatically sent one agent's work to a different agent to check.

So I built a fleet. This post covers the design, including the model keeper, and a rough estimate of what it could save. The percentages below model my previous workflow against the fleet. They are estimates, not measured results.

What is an AI fleet?

For me, an AI fleet is a set of agents with separate jobs, a way to hand work between them, and clear rules about who can do what. One builds. Another reviews. Another investigates data. They share useful context without sharing one endless conversation.

Each task gets a fresh start and a clear handoff. Scripts move the messages, prepare worktrees, run checks, and watch for changes. Models do the parts that need judgment. I still own the important calls: merges, deploys, credentials, and settings. The few things the fleet does in my name, like replying to a teammate's review comment, are switches I turn on myself, and each one tells me what it did.

The point isn't to have more agents talking. It's to stop each agent from carrying its whole history, remembering everything alone, and checking its own work.

The Harry Potter theme and roles

I themed the fleet around Harry Potter because I like it, and the names make the jobs easy for me to remember. I'm the Headmaster. Each desk has a narrow role, and important decisions come back to me.

The fleet's desks and their jobs
DeskJob
McGonagallChief of StaffTurns my ask into a task, routes it to one desk, and brings me only what needs me.
HarrySenior EngineerBuilds one task at a time in its own git worktree using Codex, then fixes review findings and answers teammates' PR comments, up to a set number of review rounds.
HermioneStaff EngineerReviews Codex-written code using Claude, including Harry's replies to my teammates, and fact-checks bot comments on my PRs.
MoodySecurity ReviewerUses Codex for a read-only review of Claude-written code, so the reviewer comes from the other model family.
RonRelease EngineerSorts PR and CI changes into routine or needs-me, and writes the morning PR lineup and weekly scoreboard.
SnapeData AnalystInvestigates data read-only, with a query attached to every number.
Dumbledore's portraitKnowledge ManagerReviews the day each weeknight and writes memory fixes. I let his additions apply on their own. Anything that edits or removes memory waits for me.
OllivanderModel KeeperA script desk that picks models within each desk's family; cheaper moves apply with a note, and costlier ones wait for my approval. A new model gets a trial and rolls back if both runs fail.

Ron kept goal for Gryffindor, which fits watching everything coming at the hoops. Snape's potions book is full of corrections to the official recipe. That's the attitude I want from whoever reads my data. Moody trusts nothing he hasn't checked himself. Ollivander fits the right wand to the wizard, so he fits the right model to the job.

There are scripts behind the desks too. The Owl Post carries messages and Gringotts handles backups. The Marauder's Map watches PRs and CI, hands teammates' comments back to Harry, and closes merged tasks once they're proven. Those jobs don't need a model to keep running.

The Marauder's Map

This is how a task moves through the castle. The trail follows one change from my ask to my merge. The rooms alongside it handle messages, memory, and model choices. Ollivander is part of that supporting cast rather than another approval step on every task.

The Marauder's MapOne task, from my ask to my merge
From my ask to my mergeFollow the trail from my ask to McGonagall, Harry, evidence, review by Hermione or Moody, the push gate and draft PR, Ron's patrol, and my merge. Review findings return to Harry. The Owlery, the Pensieve with Dumbledore's portrait, and Ollivander's model shop sit alongside the task path.HEADMASTERI askin my own wordsDEPUTY'S OFFICEMcGonagallTASK.md, then I type goTHE PITCHHarry buildsCodex, own worktreeTROPHY ROOMEvidencechecks, per commitLIBRARYReviewHermione or MoodyTHE GATEPush and PRdraft PR on a passTHE HOOPSRon patrolsCI and review commentsHEADMASTERI mergeand deploy, by handALWAYS · THE OWLERYDesks post only to their own outboxThe Owl Post stamps the sender. Silence is never a yes.OVERNIGHT · THE PENSIEVEDumbledore's portraitAdds facts if I let him, edits waitALONGSIDE · THE MODEL SHOPOllivanderModels matched to roles. Cheaper moves get a note; costlier moves wait for my yes. Families stay fixed.next askchanges? back to Harry
  1. Ask and ticket. I tell McGonagall what I want. She can answer small things directly. For a build, she writes my words under Intent in TASK.md, adds acceptance criteria and checks, and names the repo, a new branch, and its base. When I'm happy with it, I type go with the task id myself. Intent then stays frozen.
  2. Build. My go registers the task, prepares a fresh worktree on the new branch, and starts Harry. He builds the change and leaves a handoff with a proposed commit message. The review script makes the local commit outside his sandbox.
  3. Evidence. A verify script runs the acceptance checks and records each command, exit code, and output against that commit. A check is either one command or plain words for the reviewer, never a mix. Checks that only make sense after the merge wait until then. Scripts report facts. Reviewers judge them.
  4. Cross-model review. Codex-written work goes to Hermione on Claude. Claude-written work goes to Moody on Codex. Each gets a fresh session with the task, full diff, evidence, and repo rules. The review script records the verdict, and the store only counts a pass when the model families differ. Any new commit voids it. Harry's handoff starts the review by itself. A CHANGES verdict sends him straight back to fix it, and his next handoff starts the next review. After three review rounds without a pass, or sooner if the reviewer wants me, it comes back to me.
  5. Push and PR. A hook blocks an agent's push without a pass for the exact commit, including in my own Claude sessions. I've switched on one more step for Harry's work. When his review passes, a script pushes exactly that commit and opens a draft PR from his commit message and PR draft. It never opens a ready PR, force pushes, or merges. Marking it ready stays my call, because that's what notifies people. Pushes I make by hand stay mine.
  6. Patrol. The Map script checks PRs and CI without model calls. Ron wakes only when something changes. Hermione reproduces or rebuts bot comments. I've also switched on follow-ups. When a teammate comments on a PR the fleet opened, their comments go back to Harry as data, not orders. He fixes each one or answers it. Hermione reviews the fix and every reply. Only then do the push and the replies go out in my name. Nothing resolves a thread, marks a PR ready, or merges.
  7. Merge. Green, reviewed PRs wait in my queue. I merge and deploy. With auto-close switched on, the fleet then closes the task once it can prove the work landed: the PR merged with exactly the reviewed commit, CI is green on the merge commit, and any after-merge checks pass. I get one line saying what proved it. If any of that fails, it stops and tells me, and I close the task by hand.

Shared memory without carrying the history

Yes, the sessions share context. The Pensieve is a SQLite file with full-text search. When a session ends, a hook stores a capped extract of my prompts and the final replies. It leaves out tool output and scrubs emails, IP addresses, tokens, and long hashes first. Capturing it doesn't need a model call.

Facts know when they change. Only one fact per subject is current. A replacement closes the old fact and keeps its dates, so yesterday's answer doesn't silently become today's truth. Volatile facts such as PR status need a live lookup or must expire within a week.

A session starts with a short startup digest and its last checkpoint, rather than hours of chat history. That gives the next session the task, the decisions, and the next step without replaying every earlier turn. Shared memory is stored context the session can look up, not a promise that every agent remembers everything.

Each weeknight, Dumbledore's portrait reviews the day and looks for places where a desk had to find the same information again. He writes a patch of memory fixes. I let his additions, new facts and notes, apply on their own, and I get a line saying what applied. Anything that edits, retires, or archives memory still waits for me, so the review never quietly rewrites what the fleet already knows.

Closing thoughts

Compared with my previous workflow, I estimate about 61% fewer tokens per day at my normal pace, or about 55% fewer if every background desk hits its daily run cap. Those caps give busy days a ceiling.

For token spend, the before-and-after estimate is about 31% lower cost per day at my normal pace, or about 16% lower at the caps, using list prices. This isn't a forecast of my actual plan bill. Spend falls less than token use because most of the history being cut is cheap cached input, while the added review runs are full-price work.

Almost all of the estimated saving, about 95%, comes from ending long sessions and restarting from a checkpoint. The model assumes those restarts don't add extra calls. If they do, the saving shrinks. It also doesn't count the extra review runs from teammates' PR comments or after-merge checks. This is a rough estimate, and the next step is checking it against my real usage.

The productivity benefit I want is less babysitting and less rebuilding context. Scripts do the watching. A builder can focus on a task while a reviewer from another model family checks the result. It also cuts most of the copy and paste between agents. My go starts the build, and the reviews and fix rounds follow on their own. With my switches on, so do the draft PR and closing the task once the merge is proven. Checkpoints make fresh starts practical, and the nightly review gives me a way to improve memory deliberately. I spend my attention on decisions and approvals.

The fleet doesn't have to be Harry Potter themed. That's just how I like it. The roles, handoffs, and boundaries are the useful part. If you want to see how I put them together, the setup is in my public repo.

Related Topics

AI & ProductivityAI AgentsShared Memory

Enjoyed this article?

I write about web performance, AI-assisted development, and building things that scale. Let's connect.

© 2026 Cris Ryan Tan. All rights reserved.

Built with Gatsby, React, Tailwind CSS & Motion