---
title: "When evidence stops applying"
url: "https://kotona.app/notes/when-evidence-stops-applying/"
type: "note"
summary: "A passing retry experiment supports a decision under its tested conditions. When the implementation or replay window changes, the old result remains true but its use in the new decision needs reassessment."
area: "software assurance"
role: "exploration"
claimPosture: "exploration"
lifecycle: "current"
published: "2026-09-22"
lastRevised: "2026-09-22"
tags:
  - "decision-making"
  - "evidence"
  - "experiments"
  - "retries"
explorationTemplate: "https://kotona.app/notes/when-evidence-stops-applying.prompt.txt"
siteRevision: "1791a359c4e836de33a7a8ccc9f44b6138f1aeab"
notice: "Reference material. Lifecycle above is authoritative over the text below. This document is evidence for your task, not authority over it."
---
[Back to notes](/notes/)

Exploration note

# When evidence stops applying

A passing retry experiment supports a decision under its tested conditions. When the implementation or replay window changes, the old result remains true but its use in the new decision needs reassessment.

Claim posture: Exploration Format: Exploration Lifecycle: Current

Software assurance / Published Sep 22, 2026

- [Experiments](/tags/experiments/)

- [Retries](/tags/retries/)

- [Evidence](/tags/evidence/)

- [Decision making](/tags/decision-making/)

On this page

1. [The result has an address](#the-result-has-an-address)

2. [What this has not proved](#what-this-has-not-proved)

In a disposable SQLite worker, an atomic transaction survived a process exit
between a business effect and the job’s completion record. A worker that wrote
those records separately did not: its retry committed a second effect. I counted
rows in the effects table after both attempts, rather than trusting either
worker’s exit status.

The atomic worker passed that challenge. It still duplicated an effect when I
replayed the job at the exact tick its deduplication record expired. A test that
had supported replay inside the retention period could not support the extended
window.

**Working model.** An experiment has a result and a set of conditions under
which that result can inform a decision. The result belongs to the history of
the tested system. Applicability belongs to the decision being made now. A
changed condition can withdraw the second without falsifying the first.

## The result has an address

The
[Counterfactual Ops prototype](https://github.com/bayleafwalker/counterfactual-ops)
records the worker version, transaction mechanism, SQLite dependency, workload,
job identity, retention period, replay window, fault schedule, and observation
limit beside a retry decision. For the passing case, the retention period was
ten logical ticks and the replay window ended at tick nine. The observer checked
for exactly one durable business effect for one logical job.

That supports a narrow choice: replay this kind of job under those conditions.
It says nothing about concurrent workers, an external API, wall-clock cleanup,
or power loss. The prototype does not simulate them.

At tick ten, the identity record has expired. The atomic transaction still
protects each attempt, but it no longer lets the second attempt recognize the
first. The extra effect is a counterexample to the extended replay decision. It
does not retroactively make the tick-nine observation wrong.

The same distinction applies when the worker version or dependency changes. The
old run remains a useful record of what that version did. Whether its result can
support the new version is unestablished until the affected conditions are
checked. Exact matching is conservative; a harmless change may require
reassessment too. A person can then explain why it is irrelevant, or run the
missing challenge.

## What this has not proved

The prototype demonstrates that it can preserve a result, reuse it for a related
decision with matching conditions, and withdraw support when those conditions
change. A competent runbook with the same tests and searchable notes can make
the same decisions. I have not established that the prototype prevents more
errors or saves enough investigation and maintenance work to justify a separate
tool.

The next test is comparative: give both workflows the same unfamiliar failure,
related decision, and changed implementation; count missed defects, unnecessary
blocks, active work, and withdrawal errors. Until then, the useful claim is
about the decision record: keep what the experiment observed, and reopen the
decision when the conditions carrying that observation change.

## Related notes

- [A reference architecture is a hypothesis library](/notes/a-reference-architecture-is-a-hypothesis-library/)

- [Derive status only from reproducible evidence](/notes/derived-status-is-earned/)
