Keeping Track of the Models
The model roster in Homunculus reads like a staffing note, every run records which model produced it, and new releases start unproven.
Homunculus runs a small fleet of Claude agents against a real portfolio. Eight at last count — a desk manager, a risk watcher, an auditor, a pod trader that's allowed to stage real trades — plus the strategy skills underneath them. All of it rides on models that get replaced every few months. For a long time I handled that the way everyone does: pick a model once, stick it in an env var, upgrade whenever.
The trouble is that nothing breaks. A new model lands, the agents keep running, the reports keep arriving. Everything just behaves a little differently than the thing I evaluated, and with one global constant I couldn't say which decisions came from which brain. Once a couple of these agents could stage real trades, that stopped being an acceptable answer.
A roster, not a spec sheet
I stopped comparing benchmarks. A context window doesn't tell me whether a model should be driving a watcher that wakes every ten minutes. So the model list in Homunculus reads like a staffing note — each entry says what the model is for and what it isn't for.
export const AGENT_MODELS = [
{ id: 'claude-opus-5', label: 'OPUS 5',
note: 'Deepest reasoning, slowest, heaviest on your allowance.
For research and reviews — not for a watcher on a short interval.' },
{ id: 'claude-sonnet-5', label: 'SONNET 5',
note: 'Balanced. The sane default for an agent that has to reason
about the book but runs often.' },
{ id: 'claude-haiku-4-5-20251001', label: 'HAIKU 4.5',
note: 'Fastest and lightest. For high-frequency watchers whose job
is to notice a condition and escalate.' },
]Models get picked per agent, against the shape of the job. The Warden handles desk operations and wakes once a day, so it runs Haiku. The Steward reviews the whole strategy portfolio and gets Opus. Nobody puts their deepest thinker on a ten-minute polling loop, but that's exactly what a global setting does the day you flip it.
Per-agent settings also change how upgrades happen. A new model goes onto one advisory agent whose output I was reading anyway, sits there for a week, and either earns the next desk or doesn't. With a single MODEL constant, upgrading meant changing the researcher, the risk watcher, and the trader in the same commit — which in practice meant never upgrading.
Write down what actually ran
The setting only tells me what the model is now. It can't tell me what it was three weeks ago, when the fleet made a call I'm now squinting at. So every run records the model it resolved at start:
/** Model this run actually used, resolved at start ('' = server default). Recorded per
* run because the setting can change between runs, and comparing an agent's output
* across models is meaningless if you cannot tell which one produced which. */
model?: stringThis is the cheapest column in the whole system and the one I'd keep if I had to throw out the rest. Reports from the fleet get compared against each other constantly — did the Steward get sharper or did I just move it to a bigger model? Without the per-run record that question has no answer at all.
New models start unproven
The newest entry on the roster is Fable 5, and its note is pure caution: available on this subscription, no desk history with it yet, treat anything it produces as unproven until you have read a few runs. I wrote that note to myself, because six weeks from now I won't remember which models I've actually watched work and which ones I just assumed were fine.
Trust doesn't transfer between model versions. It starts over with every release.
There's one escape hatch: the validator accepts anything shaped like a Claude model id, not just the roster entries. When something new ships I can pin it by hand the same day and let it audition, without a code change standing between me and trying it.
Watch the meter
Every run also records what it consumed — fresh input tokens versus cache reads, how full the context got against that model's actual window, and whether it auto-compacted. Compaction is the number I check first. It means the context overflowed and got summarized mid-run, so the agent kept working but lost detail on the way, and its report is thinner than it reads.
Cost lands in the same record. On a subscription it isn't a bill, so I read it as allowance consumed — the number that catches a chatty watcher on the wrong model quietly eating the budget the research agent needs.
What's talking to Claude right now
The newest piece is a plain list: every live Claude session in the app — agent runs, chat turns, strategy skills, the background monitor — with its model, how long it's been going, and a stop button. Six different things in Homunculus can hold a session, and until this week nothing could even enumerate them; a wedged agent meant restarting the server. It's the least interesting screen in the app and I check it more than any other.
None of this is sophisticated. It's a list with opinions in it, one extra column on the run record, and a kill switch. But the models are going to keep churning whether I track them or not, and the difference between those two worlds is whether a model change is something I did on purpose — or something I reconstruct weeks later from a trade I don't recognize.
← Back to all posts