SAFETY · THEORY

The theory underneath, and what it lost

The research lines a harness borrows its vocabulary from, reported by what each one lost · Macheng Shen × agent · 2026-09-07

A leased action, a two-vote wrapper, a stop epoch: the mechanisms this section describes, some of them running code and some still only specified. None of them started as harness design. They are downstream of a longer research program on what information is, how it moves, and what a system forfeits when it forgets, and the harness borrowed the program's vocabulary along with, in one place, its failures. Each line below gets one paragraph on its claim, one sentence on what was falsified or remains unearned, a cognitive state, and a link. Credibility travels with what a line lost, not with what it hopes.

The spine: information dynamics and reachability

The program's spine is a claim about information itself: what matters is not a static representation but the operators that transform it and what it costs a system to forget. Aging, cancer, and a stopped agent fleet are read as instances of one integrity failure — future-reachability destroyed rather than preserved. It is the umbrella the rest of this page sits under, stated at the level of a research program, not a tested result. Front door: Living Information System and /start-here.html. State: speculative.

Multimodal representational irreducibility — a falsification, reported plainly

The claim: an embedding gap between modalities is not an optimisation failure but a structural obstruction under lossy projection — two channels projecting one latent, with more constraints than degrees of freedom, leaves a generically nonzero residual. Synthetic experiments passed cleanly: a threshold exactly where predicted, contrastive objectives stopping above an analytic floor, ten times longer training not moving it.

Then the falsification, which is the reason to report this line at all: on real models the predicted scale-decay story was tested and withdrawn — residual falls with scale under co-training, the opposite of the prediction.

What survived is narrower: a granularity dose-response, coarse concepts glue and fine ones do not, monotone across three model pairs over roughly an eightfold range; and a weakened version of the residual prediction on independently-trained towers. A later experiment reclassified contrastive co-training itself as a collapse loop rather than a repair loop — residual falls because the obstructed field is discarded, not resolved. Prior art on dimensional collapse is not yet cleared. The owner's call was to park the line as a handle rather than push it further; there is no public page for it yet. State: mechanism-level speculative; the falsified sub-claim is retired.

Agency and the self-boundary — a negative result, published

The headline is negative, and more credible for it: under thirty seeds, a candidate self-sealing dynamic could not be distinguished from generic gated slow variables or from simply freezing the action channel. A follow-up specification tests only what that negative result left untested — a boundary-coupling intervention, not a readout — and is explicitly conditional, not yet runnable. It makes no claim about phenomenal experience; that limit is part of the spec, not an omission here. Published as learning-the-self-boundary.md and its stage log. State: speculative, with a published negative result.

Memory as dynamics, not retrieval

Long-term memory is treated as an operating contract, not a vector store: a claim carries provenance, a validity interval, and an explicit conflict state, and it closes only through action — a receipt, a verification step, a measured change in behaviour. This is the direct ancestor of two sentences repeated across the mechanism and principle pages: observation is not authority, and success cannot be self-issued. Both are this position applied to one claim instead of a whole belief store. Field guide: /long-term-memory/. State: speculative.

Credit transport

Credit transport asks how reward should attribute across a graph of causes rather than a chain. An earlier headline result here — a discounted holonomy invariant — was itself withdrawn and replaced with a weighted residual against the image of the discounted incidence operator, a correction the site already carries in full. Unsettled: an internal adversarial review returned a not-yet, reframe verdict on whether the resulting credit-curvature quantity is actionable, so the line is active and partially walked back rather than closed. Full argument: discounted-credit-is-a-cokernel.md. State: speculative.

Coordination structures

A polity, a firm, a market, a commons, and a protocol are read here as five instances of one object — an architecture for scheduling scarce resources under distributed, private information, with capitalism and socialism read as two algorithms inside the same formal space rather than opposed foundations. It is where the principle layer's language of one root of authority and separate classes of principal comes from. Essay: coordination-structures.md. State: speculative.

Verification needs well-typed claims

The argument: the binding constraint on verification is not cost but whether the verifier can parse the claim at all. The measured result carries this directly — a prose-plus-regex blocker gate recalled five of ten false "waiting on the owner" parks, while a minimal typed receipt schema covering the same ground had no equivalent blind spot, not from trying harder but because a malformed typed claim has nowhere to hide the way a plausible sentence does. Full note: abundant-verification-needs-abundant-claims.html. State: speculative.

One convention, two uses

Every claim on this site, this page included, carries exactly one of three words: survived, speculative, retired. That is not decoration, and it is not unique to theory — it is the same discipline as the running/specified badge on the mechanism pages, applied to a belief instead of a piece of code. A mechanism marked specified has a design and no enforcement; a claim marked speculative has an argument and no stress test that has tried to kill it yet. Both badges exist for the same reason: so a reader, human or agent, can tell what is load-bearing without reconstructing its history. The falsifications above are not an embarrassment to route around — they are the same evidence, in the same currency, that the evidence page asks you to trust when it reports what still fails.