All tags

Tag

Evaluation

8entries across current project state and published notes.

Projects

Notes

The router was larger than the model

Exploratory claim · Current · Model evaluation

Jev, Claude Haiku 4.5 and keyword rules triaged 100 agent work items from ticket text, and none was reliably better. The ranking moved with the judge, the question's wording and the price of each error, so the next tests vary the router around the model.

The next prompt was only the visible error

Exploratory claim · Current · Agent workflow

The next-prompt heuristic measures what became visible at the human interface. A merged PR later filed as architecture, and a self-report that turned out to be reconstructable rather than introspective, both passed that test and still needed a different check further down.

The queue was never the hard part

Exploratory claim · Current · Agent infrastructure

The queue chassis behind ActionQThe PostgreSQL-backed queue that owns actions, sessions, claims, and outcomes. is now commodity, split across a Postgres-native library, a durable-workflow runtime, an execution control plane, and an automation platform. How much of ActionQ survives once that chassis is removed is the open spike.

Measure the diagnosis, not only the transcript

Archival claim · Disproven · Model evaluation

OutctlA tool that captures a command's full output outside the agent's context and returns a bounded projection the agent can query later for the omitted evidence. cut model-visible Kubernetes output by about 84%, while pair-level cost and diagnostic quality remained unresolved. The useful result was a four-part evaluation model: mechanism, quality, economics, and authority.

Judge agents by the next prompt

Guiding claim · Superseded · Agent workflow

New agents carry more work from intent to evidence, but that changes rather than removes the supervision problem. The next prompt reveals whether the result needs repair or is ready to extend.