All tags

Tag

Agents

46entries across current project state and published notes.

Projects

Notes

The router was larger than the model

Exploratory claim · Current · Model evaluation

Jev, Claude Haiku 4.5 and keyword rules triaged 100 agent work items from ticket text, and none was reliably better. The ranking moved with the judge, the question's wording and the price of each error, so the next tests vary the router around the model.

The workshop disappeared. The hotspots survived.

Exploratory claim · Current · Organizational systems

An EventStormingAlberto Brandolini's workshop method for modeling a domain on colored sticky notes -- commands, domain events, actors, views, and hotspots -- normally run with a room of stakeholders around a paper timeline. pass run solo, over agent session logs instead of a stakeholder workshop, still produced six ranked findings. The notation survived; the workshop did not; the value was in clustering the absences.

Governance for a team of one

Exploratory claim · Current · Organizational systems

A solo operator's tooling keeps adopting controls that look stolen from an organization. The controls are not for headcount; they are for coordination between actors who share neither memory nor judgment, which describes one operator across time as much as it describes a team.

The human is in the slow loop

Exploratory claim · Current · Agent workflow

A prior "human perpendicular to the loop" claim only describes one interface. The fuller model is two coupled loops at different speeds, with a settlement boundary, rather than an empirical/normative split, deciding when the slow loop must be consulted.

The next prompt was only the visible error

Exploratory claim · Current · Agent workflow

The next-prompt heuristic measures what became visible at the human interface. A merged PR later filed as architecture, and a self-report that turned out to be reconstructable rather than introspective, both passed that test and still needed a different check further down.

The queue was never the hard part

Exploratory claim · Current · Agent infrastructure

The queue chassis behind ActionQThe PostgreSQL-backed queue that owns actions, sessions, claims, and outcomes. is now commodity, split across a Postgres-native library, a durable-workflow runtime, an execution control plane, and an automation platform. How much of ActionQ survives once that chassis is removed is the open spike.

A platform capability does not exist all at once

Exploratory claim · Current · Agent infrastructure

OutctlA tool that captures a command's full output outside the agent's context and returns a bounded projection the agent can query later for the omitted evidence. began as an answer to tool results flooding model context. Programmatic tool calling now moves that work into the harness, invalidating the original product boundary while leaving a narrower durable-evidence question open.

Measure the diagnosis, not only the transcript

Archival claim · Disproven · Model evaluation

OutctlA tool that captures a command's full output outside the agent's context and returns a bounded projection the agent can query later for the omitted evidence. cut model-visible Kubernetes output by about 84%, while pair-level cost and diagnostic quality remained unresolved. The useful result was a four-part evaluation model: mechanism, quality, economics, and authority.

Why I publish explore prompts

Guiding claim · Current · Agent workflow

A note explains what survived editing. An explorePrompt is a separate, deliberately post-hoc artifact that explains where the investigation should continue, whether for a reader, their agent, or a later version of me.

The second operator is the test

Exploratory claim · Superseded · Agent workflow

sprintctlThe CLI and schema that own sprint work, dependencies, reservations, and handoffs. and actionqThe PostgreSQL-backed queue that owns actions, sessions, claims, and outcomes. enforce useful execution discipline, but five current invariants rely on one operator remaining the only source of identity, authority, and audit judgment.

Judge agents by the next prompt

Guiding claim · Superseded · Agent workflow

New agents carry more work from intent to evidence, but that changes rather than removes the supervision problem. The next prompt reveals whether the result needs repair or is ready to extend.

The ref nobody adds

Guiding claim · Current · Agent workflow

The sprint database held 831 work items and seven doc refs, six of them the same document. The reference command existed the whole time. The missing piece was never a tool — it was placement on the path the agent actually walks.

Subprocess, not service

Guiding claim · Current · Agent workflow

actionq-dispatcherRetired (2026-08-20 tombstone release). It formerly created a bounded workspace, invoked a worker, and recorded the result; the dispatch harness now constructs the workspace and invokes product-native workers directly. runs agent work as subprocesses inside a per-invocation coordinator. The rejected alternative was a long-running agent service, and the rejection is most of the design.

The aftertaste of resolution

Guiding claim · Current · Agent workflow

Conversational refinement produces the feeling of having thought clearly, which is not the same thing as having thought clearly. The fix was not discipline. It was writing the stop rule into the counterpart.

The coordinator never touches the repo

Guiding claim · Current · Agent workflow

Sprint work is split into plan, build, and review dispatches, and the orchestrating session is structurally barred from editing deliverables. The bar is the point.

The missing layer is binding, not intelligence

Archival claim · Superseded · Agent workflow

The useful workflow layer for solo agent work is not a giant orchestration stack. It is a local-first system that remembers claims, checkpoints, routing choices, and promotion boundaries without pretending one operator needs a whole platform team.