experiments.swm.cc
All experiments

Thrawn

Active

"Plan deeply, execute in parallel."

Started 5 August 2026 GitHub
Amended 27 September 2026
agentsclaude-codeorchestrationpythonherdr

The Thrawn agentic development workflow: intake, plan, execute in parallel, integrate, verify, ship

Who is Thrawn?

Grand Admiral Thrawn (Mitth’raw’nuruodo) is the Star Wars villain from Timothy Zahn’s Heir to the Empire, later canonised in Rebels and Ahsoka. He is the only Imperial commander worth fearing because he doesn’t rely on brute force. He studies his enemy first, famously through their art, until he understands how they think. Then he commits to a precise battle plan and delegates the execution to his fleet. He wins through preparation and orchestration, not firepower.

That’s the pitch for this tool. It studies the repository read-only before committing to anything, writes a battle plan, delegates the work to a fleet of agents, and nothing ships without the admiral’s sign-off. Naming it after a tactician who occasionally loses spectacularly when his subordinates improvise is, I admit, part of the joke.

What this experiment is

This is an experiment in changing my workflow to be more agentic. Not autocomplete, not pair-programming, but handing an entire ticket to a system and judging what comes back. Thrawn is the orchestrator: give it one or more GitHub or GitLab tickets, or just a markdown brief, and it:

  1. Deep-thinks a plan with a strong model that explores the repo read-only
  2. Splits the work into parallel tasks, each routed to the right model for its complexity (opus for design work, sonnet for ordinary implementation, haiku for mechanical edits, codex, pi or a local model where they fit)
  3. Spawns one agent per task, each in an isolated git worktree
  4. Merges the task branches, hands conflicts to an integrator agent, and runs the repo’s real checks
  5. Gates shipping behind a one-time code. Nothing is pushed until I’ve seen the green board and typed it

There is also a lightweight swarm mode (thrawn swarm 36 37 38 39): no planner, no integrator, just one worktree and one agent per issue with a human as the orchestrator. It has turned out to be the workhorse, and it is what made the runner trial below possible.

It is tightly coupled to my herdr setup, the terminal orchestration layer that gives every project a space and every agent a home. When thrawn spawns tasks, each task gets its own pane, so I can click through and watch exactly what each agent is doing rather than trusting a black box: every file it reads, every command it runs, every excuse it makes.

herdr running my projects: spaces down the left, agents grouped below, and a tab per session across the top

Where the experiment has been

The initial version of this page is archived in full, including the original walkthrough and a deliberately harsh self-assessment from three days in (“a parallel-agent orchestrator whose median run spawns one agent is a very expensive way to run claude -p”). That assessment produced the width gate, the approval gate and the verdict ledger, and the first essay, Your Architecture Is the Bottleneck, Not the Model, came out of what the ledger said next: roughly half my tickets don’t decompose at all, and the constraint is the shape of the codebase rather than the model.

What has happened since: the runner economics trial

A friend warned that headless Claude workers “will never benefit from caching” and would blow the bank, and recommended pi driving a GPT model on subscription instead. That became issue #8, an analysis with an experiment attached, which was then ticketed out properly and run on 28 August. The write-up is Execution Stopped Being the Bottleneck; the numbers live in the trial document.

The short version. Before the trial could run, the machinery it needed was built as its own tickets: a real pi runner (#10), pi event parsing for the activity ticker and panes (#11) and per-task token usage recorded in state.json (#12). Then the same five issues on rails_love_letter were dispatched twice with thrawn swarm, once per runner, and every branch went up as a PR so CI could referee: codex arms #58, #59, #60, #61 and #62, pi arms #63 to #67.

Both harnesses shipped credible work on all five issues. codex ran the checks unprompted and went five for five green; pi skipped linting on two branches but produced the single best branch of the ten and was the only harness whose usage thrawn could record automatically. 85 percent of pi’s token volume turned out to be server-side cache reads, which is the original warning dissolving on contact. Four codex arms and one pi arm were merged, all five issues closed, and the project’s entire game engine followed through the same machinery the next morning.

And then the platform ate the roadmap

In late September I audited the whole setup against what Claude Code now ships natively, and closed half of thrawn’s open issues in a day. Event-driven dispatch turned out to be scheduled routines with GitHub triggers. Intake grooming is a @claude mention in an Action. The pre-ship review stage is /code-review plus a persona file. Twelve open issues became two, and the full reckoning is The Platform Ate My Roadmap. What survived is the moat: routing across subscription pools (which API proxies structurally cannot do, because OAuth quota never touches a proxy), the ship gate and the panes.

The same pass re-tiered the models. Planning dropped from the top-tier model to opus, with the big brain one line of per-repo config away for tickets that earn it. Recon surveys went to haiku. A sonnet runner landed as the default mid-tier so opus only sees work that needs judgement. And the persona system that shipped inert in August got its first character: the skeptic, whose entire personality is refusing to believe a green board.

Done

  • Runner trial phase 1, scored and merged (#14)
  • pi as a first-class runner with readable panes and usage capture (#10, #11, #12)
  • Batch dispatch: thrawn 42 43 45 becomes one run planned together (#3)
  • Warm retries: same-runner retries resume the failed attempt’s session instead of starting cold (#13)
  • Abort keeps the evidence: per-task patches and head commits are snapshotted before branches are deleted, a lesson learned via git fsck (#15)
  • Integration liveness on the board, so a healthy run and a hung one look different (#2)
  • Executor and swarm prompts run the repo’s checks before committing, so pi’s lint discipline from phase 1 is fixed structurally
  • Swarm workspace detection from outside a herdr pane, by matching the repo path against pane working directories (#5)

Doing

  • Phase 2 of the runner trial, live since 27 September: pi and codex sit in the planner’s rotation for a fortnight of ordinary ungroomed work, the usage ledger records every task, and the routing verdict gets written from those numbers around 11 October (#14)
  • Interactive agent sessions (#7), the one feature idea left standing after the audit, deliberately: it leans into the panes instead of competing with native orchestration

Status

Active, and deliberately smaller than it was in August. The ledger trial continues with per-task token and cost data feeding it, and phase 2’s verdict lands around 11 October, whichever way it goes. The thesis has sharpened three times now: from “which model” to “which architecture” to “who grooms the tickets and who reviews the branches”, and September added the sharpest version: which parts of this tool can the platform never eat. The answer so far is other people’s quota, my own paranoia and my own screen.