Runner trial: codex vs pi on five real tickets
A friend warned that headless Claude workers would blow the bank and
recommended pi driving a GPT model on subscription instead. That became
an analysis (#8),
then machinery (a real pi runner, event parsing for the panes, per-task
usage capture), then a trial: the same five issues on
rails_love_letter
dispatched twice with thrawn swarm, once per runner, every branch put
up as a PR so CI could referee.
Both harnesses shipped credible work on all five. codex ran the checks unprompted and went five for five green. pi skipped linting twice but produced the single best branch of the ten, and 85 percent of its token volume turned out to be server-side cache reads, which is the original warning dissolving on contact. Four codex arms and one pi arm merged.
The write-up is Execution Stopped Being the Bottleneck; the raw scores live in the trial document.