A working family of sprint, queue, knowledge, audit, and cockpit tools now composes into one served layer, without giving up explicit state ownership or machine-local execution.
Jev, Claude Haiku 4.5 and keyword rules triaged 100 agent work items from ticket text, and none was reliably better. The ranking moved with the judge, the question's wording and the price of each error, so the next tests vary the router around the model.
A hobby game AI is a cheap laboratory for deciding which decisions should stay with an LLM, which should become learned, and which should become ordinary code, settled by replayable contests rather than diagrams.
Exploratory claim · Current · Agent infrastructure
The coordination services behind my agent work kept running after the agents moved into hosted runtimes. Because those runtimes could not reach them, the record became partial without producing an error.
Exploratory claim · Current · Agent infrastructure
Agent context is a dependency of each run: bind shared instructions to concrete versions in a per-run record, and because no linker composes prose, record every conflict outcome and refuse what cannot be classified.
Long-running agent work should continue through a validated handoff and an addressable predecessor, rather than depending on a compacted conversation as its only memory.
Exploratory claim · Current · Agent infrastructure
Model allowance, reset time, latency, cost, and concurrency should influence where agent work runs without becoming durable rules that redefine the work.
A dated map of agent-system tooling by the responsibility it primarily owns, beginning with the Markdown-and-skills setup that is enough until a specific failure earns another layer.
One session upgrades a control-plane node while another watches the API flap, unable to tell failure from maintenance. Instead of a new cluster-wide lock, a wrapper-derived maintenance notice should let the next Talos upgrade be correlated — if it cannot, the hypothesis fails.
Exploratory claim · Current · Organizational systems
An EventStormingAlberto Brandolini's workshop method for modeling a domain on colored sticky notes -- commands, domain events, actors, views, and hotspots -- normally run with a room of stakeholders around a paper timeline. pass run solo, over agent session logs instead of a stakeholder workshop, still produced six ranked findings. The notation survived; the workshop did not; the value was in clustering the absences.
Exploratory claim · Current · Organizational systems
A solo operator's tooling keeps adopting controls that look stolen from an organization. The controls are not for headcount; they are for coordination between actors who share neither memory nor judgment, which describes one operator across time as much as it describes a team.
A prior "human perpendicular to the loop" claim only describes one interface. The fuller model is two coupled loops at different speeds, with a settlement boundary, rather than an empirical/normative split, deciding when the slow loop must be consulted.
The next-prompt heuristic measures what became visible at the human interface. A merged PR later filed as architecture, and a self-report that turned out to be reconstructable rather than introspective, both passed that test and still needed a different check further down.
RAG, vector search, graph engines, and generic evaluation tooling are moving into managed platforms. The durable work is defining which evidence may count, what failure means, and when an agentic system is acceptable to ship.
Exploratory claim · Current · Agent infrastructure
The queue chassis behind ActionQThe PostgreSQL-backed queue that owns actions, sessions, claims, and outcomes. is now commodity, split across a Postgres-native library, a durable-workflow runtime, an execution control plane, and an automation platform. How much of ActionQ survives once that chassis is removed is the open spike.
Exploratory claim · Current · Agent infrastructure
OutctlA tool that captures a command's full output outside the agent's context and returns a bounded projection the agent can query later for the omitted evidence. began as an answer to tool results flooding model context. Programmatic tool calling now moves that work into the harness, invalidating the original product boundary while leaving a narrower durable-evidence question open.
Exploratory claim · Current · Agent infrastructure
Production access exposes the limits of prompt-only agent design. A useful operational run binds a general harness to domain context, scoped authority, evidence, and completion rules.
An agent asks for a design decision mid-implementation, the operator answers at pull-request depth, and the answer gets filed as architecture. The fix is not a better escalation prompt but keeping three separate reviews from borrowing each other's authority.
OutctlA tool that captures a command's full output outside the agent's context and returns a bounded projection the agent can query later for the omitted evidence. cut model-visible Kubernetes output by about 84%, while pair-level cost and diagnostic quality remained unresolved. The useful result was a four-part evaluation model: mechanism, quality, economics, and authority.
A multi-repository project folder may assemble member worktrees, shared guidance and project context, but it must not acquire ownership of Git, backlog, execution or deployment state merely because every tool can see it.
A small isolated-completion pilot asked whether a model can report influences on its own output. External-reconstruction controls qualitatively reproduced the effects, so the useful residue was operational, not introspective.
Direct agent operation removes a universal deployment handoff, so authority, evidence, and reconciliation must bind to each consequential action instead of to its location.
A reading map for building AI systems that can show what they were asked to do, what they did, why a result should be trusted, and how failures are contained. It starts with established assurance practice and ends with a dated watch list of open work.
Chronology is one projection of this site's notes, not the structure underneath them. The corpus is the artifact of record; pages, feeds, and prompts built from it are derived realizations that should stay regenerable rather than independently maintained.
The assurance questions behind my recent agent-workflow notes are established engineering problems. The useful work is learning their vocabulary and testing their methods at small scale.
A note explains what survived editing. An explorePrompt is a separate, deliberately post-hoc artifact that explains where the investigation should continue, whether for a reader, their agent, or a later version of me.
Continuing assurance can move away from an exact generated output only when retained inputs and independent checks can produce another acceptable result; retention remains a separate decision.
Working model: a production-critical agent pipeline becomes an institutional capability only when a second person can reconstruct, operate, change, and recover it from recorded substrate and evidence.
Exploratory claim · Current · Organizational systems
Agent-ready work keeps the reason, permission, attempted action, observed result, and correction connected while the work happens instead of reconstructing them later.
A durable agent application keeps permissions, records, checks, and recovery outside the model, so changing models does not also replace the system's memory or rules.
Agents can bypass the physical handoff between execution and runtime, so authorization, evidence, and reconciliation have to attach to each consequential action instead.
A devbox can bind identity, tools, network reach, and session evidence, but it should remain a replaceable access cell rather than becoming the organizational authority.
Exploratory claim · Superseded · Organizational systems
Routine junior work once paid for useful production and professional formation at the same time; agent automation separates those goods and leaves succession needing an explicit operating model.
AI makes synthesis cheap enough that the prose has little scarcity. The human remainder is the changed judgment and an attributable position that can age badly in public.
AI weakens the old correlation between producing a plausible patch and understanding a project, so OSS contribution shifts toward evidence, verification, standing, and durable ownership.
sprintctlThe CLI and schema that own sprint work, dependencies, reservations, and handoffs. and actionqThe PostgreSQL-backed queue that owns actions, sessions, claims, and outcomes. enforce useful execution discipline, but five current invariants rely on one operator remaining the only source of identity, authority, and audit judgment.
Work-management systems can delegate and coding agents can execute, but authority, attempts, evidence, and acceptance still have no obvious neutral owner between them.
New agents carry more work from intent to evidence, but that changes rather than removes the supervision problem. The next prompt reveals whether the result needs repair or is ready to extend.
The sprint database held 831 work items and seven doc refs, six of them the same document. The reference command existed the whole time. The missing piece was never a tool — it was placement on the path the agent actually walks.
actionq-dispatcherRetired (2026-08-20 tombstone release). It formerly created a bounded workspace, invoked a worker, and recorded the result; the dispatch harness now constructs the workspace and invokes product-native workers directly. runs agent work as subprocesses inside a per-invocation coordinator. The rejected alternative was a long-running agent service, and the rejection is most of the design.
Conversational refinement produces the feeling of having thought clearly, which is not the same thing as having thought clearly. The fix was not discipline. It was writing the stop rule into the counterpart.
Sprint work is split into plan, build, and review dispatches, and the orchestrating session is structurally barred from editing deliverables. The bar is the point.
The useful workflow layer for solo agent work is not a giant orchestration stack. It is a local-first system that remembers claims, checkpoints, routing choices, and promotion boundaries without pretending one operator needs a whole platform team.