Thrawn
Active"Plan deeply, execute in parallel."

Who is Thrawn?
Grand Admiral Thrawn (Mitth’raw’nuruodo) is the Star Wars villain from Timothy Zahn’s Heir to the Empire, later canonised in Rebels and Ahsoka. He is the only Imperial commander worth fearing because he doesn’t rely on brute force. He studies his enemy first, famously through their art, until he understands how they think. Then he commits to a precise battle plan and delegates the execution to his fleet. He wins through preparation and orchestration, not firepower.
That’s the pitch for this tool. It studies the repository read-only before committing to anything, writes a battle plan, delegates the work to a fleet of agents, and nothing ships without the admiral’s sign-off. Naming it after a tactician who occasionally loses spectacularly when his subordinates improvise is, I admit, part of the joke.
What this experiment is
This is an experiment in changing my workflow to be more agentic. Not autocomplete, not pair-programming, but handing an entire ticket to a system and judging what comes back. Thrawn is the orchestrator: give it one or more GitHub or GitLab tickets, or just a markdown brief, and it:
- Deep-thinks a plan with a strong model that explores the repo read-only
- Splits the work into parallel tasks, each routed to the right model for its complexity (opus for design work, sonnet for ordinary implementation, haiku for mechanical edits, codex, pi or a local model where they fit)
- Spawns one agent per task, each in an isolated git worktree
- Merges the task branches, hands conflicts to an integrator agent, and runs the repo’s real checks
- Gates shipping behind a one-time code. Nothing is pushed until I’ve seen the green board and typed it
There is also a lightweight swarm mode (thrawn swarm 36 37 38 39): no
planner, no integrator, just one worktree and one agent per issue with a
human as the orchestrator. It has turned out to be the workhorse, and it is
what made the runner trial below possible.
It is tightly coupled to my herdr setup, the terminal orchestration layer that gives every project a space and every agent a home. When thrawn spawns tasks, each task gets its own pane, so I can click through and watch exactly what each agent is doing rather than trusting a black box: every file it reads, every command it runs, every excuse it makes.

Where the experiment has been
The initial version of this page is
archived in full, including the original walkthrough and a deliberately
harsh self-assessment from three days in (“a parallel-agent orchestrator
whose median run spawns one agent is a very expensive way to run
claude -p”). That assessment produced the width gate, the approval gate
and the verdict ledger, and the first essay,
Your Architecture Is the Bottleneck, Not the Model,
came out of what the ledger said next: roughly half my tickets don’t
decompose at all, and the constraint is the shape of the codebase rather
than the model.
What has happened since: the runner economics trial
A friend warned that headless Claude workers “will never benefit from caching” and would blow the bank, and recommended pi driving a GPT model on subscription instead. That became issue #8, an analysis with an experiment attached, which was then ticketed out properly and run on 28 August. The write-up is Execution Stopped Being the Bottleneck; the numbers live in the trial document.
The short version. Before the trial could run, the machinery it needed was
built as its own tickets: a real pi runner
(#10), pi event
parsing for the activity ticker and panes
(#11) and
per-task token usage recorded in state.json
(#12). Then the
same five issues on
rails_love_letter were
dispatched twice with thrawn swarm, once per runner, and every branch
went up as a PR so CI could referee: codex arms
#58,
#59,
#60,
#61 and
#62, pi arms
#63 to
#67.
Both harnesses shipped credible work on all five issues. codex ran the checks unprompted and went five for five green; pi skipped linting on two branches but produced the single best branch of the ten and was the only harness whose usage thrawn could record automatically. 85 percent of pi’s token volume turned out to be server-side cache reads, which is the original warning dissolving on contact. Four codex arms and one pi arm were merged, all five issues closed, and the project’s entire game engine followed through the same machinery the next morning.
And then the platform ate the roadmap
In late September I audited the whole setup against what Claude Code now
ships natively, and closed half of thrawn’s open issues in a day.
Event-driven dispatch turned out to be scheduled routines with GitHub
triggers. Intake grooming is a @claude mention in an Action. The pre-ship
review stage is /code-review plus a persona file. Twelve open issues
became two, and the full reckoning is
The Platform Ate My Roadmap.
What survived is the moat: routing across subscription pools (which API
proxies structurally cannot do, because OAuth quota never touches a proxy),
the ship gate and the panes.
The same pass re-tiered the models. Planning dropped from the top-tier model to opus, with the big brain one line of per-repo config away for tickets that earn it. Recon surveys went to haiku. A sonnet runner landed as the default mid-tier so opus only sees work that needs judgement. And the persona system that shipped inert in August got its first character: the skeptic, whose entire personality is refusing to believe a green board.
Done
- Runner trial phase 1, scored and merged (#14)
- pi as a first-class runner with readable panes and usage capture (#10, #11, #12)
- Batch dispatch:
thrawn 42 43 45becomes one run planned together (#3) - Warm retries: same-runner retries resume the failed attempt’s session instead of starting cold (#13)
- Abort keeps the evidence: per-task patches and head commits are
snapshotted before branches are deleted, a lesson learned via
git fsck(#15) - Integration liveness on the board, so a healthy run and a hung one look different (#2)
- Executor and swarm prompts run the repo’s checks before committing, so pi’s lint discipline from phase 1 is fixed structurally
- Swarm workspace detection from outside a herdr pane, by matching the repo path against pane working directories (#5)
Doing
- Phase 2 of the runner trial, live since 27 September: pi and codex sit in the planner’s rotation for a fortnight of ordinary ungroomed work, the usage ledger records every task, and the routing verdict gets written from those numbers around 11 October (#14)
- Interactive agent sessions (#7), the one feature idea left standing after the audit, deliberately: it leans into the panes instead of competing with native orchestration
Status
Active, and deliberately smaller than it was in August. The ledger trial continues with per-task token and cost data feeding it, and phase 2’s verdict lands around 11 October, whichever way it goes. The thesis has sharpened three times now: from “which model” to “which architecture” to “who grooms the tickets and who reviews the branches”, and September added the sharpest version: which parts of this tool can the platform never eat. The answer so far is other people’s quota, my own paranoia and my own screen.