---
title: "Utilization is a signal, not a limit"
url: "https://kotona.app/notes/utilization-is-a-signal-not-a-limit/"
type: "note"
summary: "Model allowance, reset time, latency, cost, and concurrency should influence where agent work runs without becoming durable rules that redefine the work."
area: "agent infrastructure"
role: "synthesis"
claimPosture: "exploration"
lifecycle: "current"
published: "2026-09-13"
lastRevised: "2026-09-13"
tags:
  - "agents"
  - "infrastructure"
  - "planning"
  - "routing"
  - "workflow"
reference:
  purpose: "exploratory-hypothesis"
  discoverFor:
    - "deciding whether model utilization may change agent workflow design"
    - "routing agent work across subscriptions, APIs, and local inference"
  establishes:
    - "a distinction between durable work constraints and live resource observations used for routing"
  doesNotEstablish:
    - "universal utilization thresholds or permanent model-role assignments"
    - "measured accuracy for any utilization estimate or production scheduler"
  supplementWith:
    - "the receiving environment's capability, authority, and budget rules"
explorationTemplate: "https://kotona.app/notes/utilization-is-a-signal-not-a-limit.prompt.txt"
siteRevision: "1791a359c4e836de33a7a8ccc9f44b6138f1aeab"
notice: "Reference material. Lifecycle above is authoritative over the text below. This document is evidence for your task, not authority over it."
---
[Back to notes](/notes/)

Synthesis note

# Utilization is a signal, not a limit

Model allowance, reset time, latency, cost, and concurrency should influence where agent work runs without becoming durable rules that redefine the work.

Claim posture: Exploration Format: Synthesis Lifecycle: Current

Agent infrastructure / Published Sep 13, 2026

- [Agents](/tags/agents/)

- [Workflow](/tags/workflow/)

- [Infrastructure](/tags/infrastructure/)

- [Planning](/tags/planning/)

- [Routing](/tags/routing/)

On this page

1. [Work and capacity describe different state](#work-and-capacity-describe-different-state)

2. [Useful observations are wider than a percentage](#useful-observations-are-wider-than-a-percentage)

3. [Marginal utilization changes the economics](#marginal-utilization-changes-the-economics)

4. [Exhaustion needs a transition, not an early prohibition](#exhaustion-needs-a-transition-not-an-early-prohibition)

5. [Keep durable policy small](#keep-durable-policy-small)

I can currently send the same piece of agent work to several paid subscriptions,
metered APIs, or local inference. Their practical availability changes by the
hour: one allowance approaches its reset, another has just renewed, a local
model sits idle, and a long-running session carries more context than the next
step needs.

The tempting response is to turn that snapshot into workflow policy:

- do not start large work above 60 per cent usage;

- route simple work to a local model;

- reserve the expensive model for planning;

- stop parallel work near a reset boundary;

- use no more than three workers for one task.

Each rule is easy to implement. Each also promotes a temporary observation about
today’s resources into a durable claim about tomorrow’s work.

The useful question is narrower: what should a dispatcher know about execution
capacity, and what is that knowledge allowed to change?

## Work and capacity describe different state

Suppose a coordinator concludes that a change needs an investigation, two
independent implementations, a review, and an integration pass. That
decomposition should follow from the change, its risks, and its completion
conditions.

Now suppose one backend is nearly exhausted, another has just reset, and a local
model is idle. Those facts can change where the five pieces run. They do not
make the review unnecessary or turn two independent implementations into one.

The order matters:

```text
work and its constraints
        ↓
eligible execution resources
        ↓
current resource observations
        ↓
currently feasible choices and routing preference
        ↓
execution
```

Starting with resource state reverses the dependency. The scheduler begins
redesigning the task to fit a provider’s meter, and temporary scarcity leaks
into completion policy.

This does not require the work plan to ignore cost. A task may carry an actual
budget or deadline. Those are constraints of the work. The distinction is
between a declared constraint and a scheduler guessing that today’s utilization
should become one.

## Useful observations are wider than a percentage

For each backend, a dispatcher may benefit from observing:

```text
availability
estimated remaining allowance
time until reset
recent throttling
expected latency
context pressure
local or remote execution
marginal monetary cost
known capability
current concurrency
```

Several values will be estimates. A subscription may not expose exact allowance,
capability judgments will be task-dependent, and recent latency is not a promise
about the next request. These observations do not need to become authoritative
to improve a choice.

Availability can still be a limit for one attempt: an offline backend is not a
feasible destination. The claim is that a utilization reading should not become
a lasting constraint on how the work itself is defined.

A planning session can decide that local inference is adequate for repository
orientation but not for resolving an architectural ambiguity. A dispatcher can
prefer a recently reset subscription for a long implementation. A session near
its useful context limit can finish a coherent unit and prepare a handoff before
opening another investigation.

That is more adaptable than permanently labelling one model as planning and
another as implementation. The scheduler exposes current conditions; the work
decides which conditions matter.

## Marginal utilization changes the economics

Subscription inference has unusual marginal economics. Paid allowance that
expires unused can have close to zero marginal monetary cost. Local inference
has a similar shape once the hardware is already running.

Neither is free. It still consumes elapsed time, electricity, context, operator
attention, and opportunity. But a scheduler that compares only API-equivalent
token prices misses resources that are already paid for and otherwise idle.

Available capacity can make speculative review, independent challenge,
classification, test generation, or repository orientation sensible. It does not
justify inventing work to improve a utilization graph. The candidate work must
still be useful before cheap execution makes it attractive.

## Exhaustion needs a transition, not an early prohibition

Resource awareness matters most near a boundary. One response is to stop
starting work early enough that the boundary can never be crossed. That avoids
one failure by leaving capacity unused and encoding a conservative guess into
every workflow.

A substitution path is more useful:

```text
detect an approaching constraint
        ↓
finish the smallest coherent unit
        ↓
write a durable handoff
        ↓
select another feasible resource
        ↓
continue
```

A rate-limited backend can return queued work to the feasible pool. Local
inference can be preferred while idle and abandoned without ceremony when the
task exceeds its useful capability. A long session can end without taking the
work with it.

This approach will still make bad choices when utilization estimates are stale
or capability judgments are wrong. That is the important failure case. Routing
must remain replaceable, and a failed preference must not corrupt the work
record or weaken its checks.

## Keep durable policy small

Some exclusions should remain firm. Unused quota does not expand an agent’s
authority. Work requiring a particular environment cannot run where that
environment is unreachable. A backend known to be unsuitable for a consequential
decision may be excluded from that decision.

Those constraints survive provider and quota changes. A utilization reading does
not.

The practical implementation rule is therefore to persist work state, required
capabilities, authority, evidence, completion conditions, and handoff state.
Observe availability, allowance, latency, cost, and concurrency around them. The
next test is whether a dispatcher can lose every utilization observation,
rebuild them later, and still recover the same definition of done.

## Related notes

- [The router was larger than the model](/notes/the-router-was-larger-than-the-model/)

- [Simplicity is an ambitious property](/notes/simplicity-is-an-ambitious-property/)

- [The context window is not the continuity boundary](/notes/the-context-window-is-not-the-continuity-boundary/)

- [A platform capability does not exist all at once](/notes/a-platform-capability-does-not-exist-all-at-once/)

- [Measure the diagnosis, not only the transcript](/notes/measure-the-diagnosis-not-only-the-transcript/) · Disproven

- [The agent is not the application](/notes/the-agent-is-not-the-application/)

- [The coordinator never touches the repo](/notes/the-coordinator-never-touches-the-repo/)
