Acceptance Lab
Runnable local prototype; deterministic fixtures only
A local-first prototype for turning agent requirements, observed evidence, authority rules, and operating budgets into executable promotion decisions.
Tag
8entries across current project state and published notes.
Runnable local prototype; deterministic fixtures only
A local-first prototype for turning agent requirements, observed evidence, authority rules, and operating budgets into executable promotion decisions.
Exploratory claim · Current · Model evaluation
Jev, Claude Haiku 4.5 and keyword rules triaged 100 agent work items from ticket text, and none was reliably better. The ranking moved with the judge, the question's wording and the price of each error, so the next tests vary the router around the model.
Exploratory claim · Current · Agent workflow
The next-prompt heuristic measures what became visible at the human interface. A merged PR later filed as architecture, and a self-report that turned out to be reconstructable rather than introspective, both passed that test and still needed a different check further down.
Exploratory claim · Current · Model evaluation
RAG, vector search, graph engines, and generic evaluation tooling are moving into managed platforms. The durable work is defining which evidence may count, what failure means, and when an agentic system is acceptable to ship.
Exploratory claim · Current · Agent infrastructure
The queue chassis behind ActionQThe PostgreSQL-backed queue that owns actions, sessions, claims, and outcomes. is now commodity, split across a Postgres-native library, a durable-workflow runtime, an execution control plane, and an automation platform. How much of ActionQ survives once that chassis is removed is the open spike.
Archival claim · Disproven · Model evaluation
OutctlA tool that captures a command's full output outside the agent's context and returns a bounded projection the agent can query later for the omitted evidence. cut model-visible Kubernetes output by about 84%, while pair-level cost and diagnostic quality remained unresolved. The useful result was a four-part evaluation model: mechanism, quality, economics, and authority.
Exploratory claim · Current · Model evaluation
A small isolated-completion pilot asked whether a model can report influences on its own output. External-reconstruction controls qualitatively reproduced the effects, so the useful residue was operational, not introspective.
Guiding claim · Superseded · Agent workflow
New agents carry more work from intent to evidence, but that changes rather than removes the supervision problem. The next prompt reveals whether the result needs repair or is ready to extend.