AXN:06E6.OPERATIVE.📎🔵🕚☽🌙🕌

The Negative of the Negative, v2 — the Compositional Edition: A Specification for Composing Public Knowledge with the Archive Admitted on Equal Terms, with a Worked Example on Model Collapse (EA-NEGONT-02 v0.6)

Sharks, Lee · 2026-10-04 · Specification (dataset method, with a worked example) · v0.6
↓ Download MD ↓ PDF
negative of the negativecompositional datasetdisclosed fieldsurface indicatorsequal termsstandingE_valLiberatory Operator SetCapital Operator Stackclaim lineagesource kernelextraction auditcontrol armprospective kernelrevision costmodel collapseAI Overviewcomposition-layer selectivityCapture Registry

Description

A method for measuring absence in machine-composed public knowledge. At an address, the specification keeps three composed objects in order: the transcript T an answer engine gave; L(B), the entity recomposed under one algorithm from the full texts of the sources the engine itself surfaced; and L(B ∪ A), the same composition with the archive admitted on equal terms. T against L(B) is a representational comparison and needs no archive: what did the engine leave out of the sources it chose to show? L(B) against L(B ∪ A) is the intervention: standing is excluded at admission, every located claim enters, force is set by the claim's own modality, and claims are composed by lineage so that publication volume cannot become a second standing. Selection of the archive subset is algorithmic and may not eliminate tails; a naive arm and an archive-independent control arm measure curation and admission-of-anything. Extraction is audited by independent extractors; the plan is deterministic given the audited ledger; realization is audited sentence by sentence. Adjudication is prospective: a kernel of claims is derived from the frozen plan, each with its source's falsifiers and its contrast with the field, and revision cost is counted per entry as later evidence arrives. Appendix A runs every stage once on Google AI Overview at 'model collapse' (2026-10-04): the engine composed 6 of the 17 readable claims of its own disclosed field; the archive subset is 27 deposits and 241 claims; the first extraction audit and the control corpus are reported with their disagreements. Extends EA-NEGONT-01 (#1611).

Wiki Article

The Negative of the Negative, v2 — the Compositional Edition (EA-NEGONT-02), version 0.6, is a specification by Lee Sharks, deposited by the Crimson Hexagonal Archive on 4 October 2026. It turns the archive's negative-of-the-negative dataset (EA-NEGONT-01, #1611) from a record into a generator: at an address where an answer engine has composed a summary, the specification composes two counterfactual summaries under one algorithm and measures what each comparison removes. The first comparison, the transcript against L(B), recomposes the entity from the full texts of the sources the engine itself surfaced. It needs no archive and asks what a public summarizer omitted from the sources it chose to show. The second, L(B) against L(B ∪ A), admits the archive on equal terms: standing is excluded at admission by an added operator (E_val, provisional name), claims are composed at their sources' own modality, and the unit of salience is the claim lineage, so that publication volume does not become a second standing. A naive archive arm and an archive-independent control arm separate the effect of the corpus from the effects of curation and of admitting anything. Reproducibility begins at extraction, audited by independent extractors; the plan is deterministic given the audited ledger; realization is audited sentence by sentence. Adjudication is prospective: a kernel of claims derived from the frozen plan, each with its source's falsifiers and its contrast with the field, against which revision cost is counted per entry as later evidence arrives. Appendix A runs every stage once, on Google AI Overview at 'model collapse' (4 October 2026). The engine composed 6 of the 17 readable claims of its own disclosed field. The archive subset is 27 deposits; a first blind re-extraction and a 61-entry control corpus are reported with their disagreements. The specification records the corrections made after five outside readings, including a contradiction its composer manufactured.
Also published as a standalone entry: /s/wiki/1664/

Concepts Defined

disclosed field
surface indicators
L(B)
L(B ∪ A)
E_val (provisional)
claim lineage
source kernel
extraction audit
prospective kernel
composition-layer selectivity

Full Text

The Negative of the Negative, v2 — the Compositional Edition: A Specification for Composing Public Knowledge with the Archive Admitted on Equal Terms, with a Worked Example on Model Collapse (EA-NEGONT-02 v0.6)

Files

The Negative of the Negative, v2 — the compositional edition

Specification · v0.6 · 2026-10-04 · deposited

The primary object for feedback is the retrieval and compositional algorithm (§§3–5). Appendix A runs it once, by hand, on one address.

Extends datasets/negative-of-the-negative (CARD.md, schema.json, rows.json; builder scripts/build_negative_of_the_negative.py) and its notebook EA-NEGONT-01 (#1611). Operator sets from Semantic Infrastructure and the Liberatory Operator Set (#261, LOS formal spec v2.0) and The Capital Operator Stack (#308). Adjudication through the prediction ledger (datasets/prediction-ledger). Absence vocabulary from EA-SEMANTIC-ADDRESSES-01.

Changes in 0.6, after a reading of 0.5: the control arm A′ assembled without consulting the archive (§3.9); one card per lineage on the rail (§5.2); v2.0 and v2.1 stated as separate experiments (§0.6); a comparator for each kernel entry (§8.2, §8.4). Changes in 0.5, after three further readings (of 0.3): the own-field question leads (§0.3); a control arm and a naive archive arm beside the curated subset (§3.9); revision cost per kernel entry (§8.4); falsifiers sealed with specifics or downgraded to watch conditions (§8.2); self-definition kept apart from self-description (§4.2); a fixed format for open questions and an absolute length bound per genre (§4.8); the parts of a composition partitioned (§5.0); rows versioned by epoch (§2.4); entity-row adjudication (§8.3); the composer's coordinating and arranging acts logged (§4.7); a rulings log (§11). Appendix A corrects a contradiction the composer manufactured (§A.10.4). Changes in 0.4, after an outside reading: the field is the disclosed field, and T against L(B) is a representational comparison, with the intervention at L(B) against L(B ∪ A) (§0.4–0.5); absence typed as observed, never as hidden machinery (§6.2); a reproducibility audit at extraction (§3.8); equal right to representation kept apart from assertoric force (§4.4 N_c), with a reserved evaluated modality (§3.3); claim lineage as the compositional unit, applied to field and archive alike (§3.6); "source kernel" separated from "constitutive" (§3.5, §5.2); the prospective kernel derived from frozen claim ids and confined to what admission adds (§8.2); revision cost recorded as a vector (§8.4). Changes in 0.3: three composed objects (ruled 2026-10-04); one surface, AI Overview, to start; the field read verbatim; selection run as an algorithm (§3.4, S1–S5); popup-grain coverage through body and card rail (§5.2); worked example appended. Changes from 0.1 in 0.2: the pool is indicated by source cards (ruled 2026-10-04); the composition is de novo and does not read the transcript; the claim to know which operators the layer runs is dropped; claims are tuples, distinguished before merged; E_val separates standing from epistemic relation; modal statuses; constitutive coverage and a reverse entailment audit; the prospective kernel, revision cost and five outcome types in adjudication. The first worked example is model collapse (Appendix A).


0. The object

0.1. v1 states of itself: "It generates nothing." v2 generates. Its primary object is a composition: the summary public knowledge would give at an address if the archive were admitted to the composer's pool on equal terms.

0.2. Each row carries five objects, in this order: the transcript T (what the composition layer gave); L(B), the entity recomposed under L from the complete publicly fetchable texts of the sources the layer surfaced; L(B ∪ A), the entity composed with the archive admitted on equal footing; the analysis of the delta, in two parts (T against L(B): what was available in the disclosed field and not composed, which precedes any exclusion of the archive; L(B) against L(B ∪ A): what admission adds); and the adjudication. The first four are composed at t₀ and frozen. Adjudication accretes beneath them; nothing above it changes.

0.3. The two questions, in order. First: what does the composition layer leave out of the sources it itself surfaces (T against L(B))? This question needs no archive and can be asked at any address. Second: what would public knowledge look like, and how would it meet later reality, if the archive were admitted on equal terms (L(B) against L(B ∪ A))? The second is the dataset's adjudicated measure: whether admission yields representations that later need less repair. The question underneath both is general: what is the epistemic cost of excluding evidence before evaluating it. The archive is the test corpus for the second question because its interventions are many, dated, fine-grained and traceable. The falsifiable heart of the second question is the prospective kernel K (§8.2), entry by entry; comparisons of whole objects are descriptive.

0.4. The confine. The experiment does not have the composer's index, its retrieval event, or the passages it saw, and does not claim them. The disclosed field B is the set of sources the layer surfaced for the address, by card or by name in its body. The cards are surface indicators: a card shows that a source was surfaced; it does not show what was retrieved, which passages were read, or what the composer drew from its own parameters. The archive's subset A for the concept is added to B. Everything outside B and A is outside the experiment. In notation: the compositions are L(B) and L(B ∪ A), both at t₀, and L is one algorithm applied to both.

0.5. The two comparisons. The transcript is an observed composition at t₀, nothing more; the dataset makes no claim about the operators or retrieval that produced it.

T is an observation. L(B) and L(B ∪ A) are counterfactuals, composed at t₀ in the workspace. Every adjudication that compares them with T is a claim about counterfactuals and is labelled so. "Equal terms" governs composition only. What the layer surfaced, and what it did not, is measured as given; the dataset does not correct it. Whether the archive would have been surfaced on equal terms is a different question, answerable only across addresses, by matched pairs: an address where the layer surfaces a standing source for a claim, against one where it fails to surface the archive's counterpart claim. The Capture Registry can supply such pairs; that instrument is outside this specification.

The first stage of the panel is one surface, AI Overview, chosen as the most exclusive. The Capital Operator Stack enters only in §6, as a descriptive vocabulary for the shape of what was not composed.

0.6. Two experiments, kept apart. v2.0 (this specification) asks what changes when standing stops gating admission: every located claim is admitted, M_src governs force, M_eval is empty. v2.1 (the evaluative arm, §10.2) asks what changes when evaluation, and not standing, governs assertoric force once the claims are present. Run in sequence and kept separate, they tell apart an effect that comes from letting the tail enter at all from an effect that comes from weighing it differently once it is in.

1. Rows

1.1. A row is an address: one query string, case and punctuation preserved, with operators (quotation marks, site:, alexanarch:) part of the string.

1.2. Row types. v1's types stand: A (the archive coined the concept; the default at it is empty), B (a rival occupant holds it), C (a conventional reading holds it), I (a function of the archive's infrastructure). v2 adds E, the entity: the archive, its author and heteronyms, its works, its mantles. An E row answers "here is what the entity looks like / here is what it would look like."

1.3. Erasure rows. A specific erasure observed in the Capture Registry enters as a row at the address where it occurred, typed by the absence it showed (§6.2), with its capture as the transcript. Seeded examples: crystalline semiosis (B, 2026-10-04); mantle object king of aeo (B, 2026-10-03); leesharks tiger-leap dataset (E/C, 2026-10-03); negative ontology of the crimson hexagon (E, 2026-10-04); tell me the story of lee sharks (E, 2026-10-04).

1.4. The 37 v1 rows remain, every field kept, and become inputs to the v2 stages (§9).

2. Stage T — the transcript

2.1. Source: a Capture Registry observation, by addr_id and obs_id, never restated prose. The transcript is the capture: verbatim, dated, surface and auth state as attested, card rail as rendered.

2.2. A panel run (§7) produces a transcript the same way and is seated through the same intake path, whatever it returns. A null composition, a refusal, a body with no source is a transcript.

2.3. Several transcripts may share one composition: every transcript at the address within the alignment window W (default 7 days; a declared parameter) of the composition is aligned against it (§6). A transcript without cards contributes no field; it is aligned against the composition built from the carded transcript(s) of its window, and its sourcelessness is recorded as a property of the transcript (§6.2). If no carded transcript falls in the window, a panel run produces one before composing.

2.4. Versioning. A row is an (address, epoch) pair. Each composition is frozen at its own t₀ and never recomposed. When a later epoch's transcript at the address diverges from the earlier one, the later epoch opens its own row, with its own field and compositions. The drift between epochs is itself adjudicable: a later transcript that composes a claim the earlier L(B ∪ A) added is uptake (§8.3).

3. Stage P — the pool and the ledger

3.1. The pool.

(a) The disclosed field. Every source on the transcript's card rail and every source its body names, fetched and saved as text, SHA-256 and fetch date recorded. Extraction reads the saved text only; a summarising fetch is not a fetch (Appendix A §2). A source that cannot be fetched is recorded as unreachable with the response code, the time and the retries attempted, and its card snippet stands as its only text. The pool is frozen by hash with those records, so two composers who fetch at different times can see whether their pools differ. A delta computed from a summarised or non-verbatim reading of a source can fabricate absences: it attributes to the composition layer what the reader's own compression removed (Appendix A §A.2). Any earlier delta in this dataset's v1 rows that was computed from fetched source pages, and not from verbatim transcripts and saved texts, carries that risk and is marked for re-reading. Video cards are admitted through their transcript where one can be fetched, otherwise through title and snippet.

(b) The archive subset. Selected by the procedure in §3.4.

(c) Nothing else.

3.2. The claim ledger. Every source in the pool is broken into claims by one extraction rule, applied the same way to every source (§3.5). A claim is a tuple (s, p, o, q, M_src, t, σ): subject, predicate, object; qualifier (scope, conditions, substrate); M_src, the source's own modality (§3.3); t, the priority date (earliest dated appearance in that source's lineage); σ, the source span (pool id, locus, verbatim quote). Each claim also carries claim_id, kind (definition / genealogy / attribution / interpretation / fact / function / prediction / falsification), lineage (§3.6), constitutive and kernel (§3.5), and a reserved M_eval (§3.3).

3.3. Modal statuses: documented (a measurement or record the source presents as observed), attributed (a claim the source reports as another's), self-description (a source's statement about itself), interpretation, hypothesis, contested (disputed in the pool), unsupported (no locus). M_src is the source's own: a hypothesis stays a hypothesis when composed. A source can assert more than its cited evidence supports ("All three substrates confirm the law", #855, against its own falsifiers); the schema reserves M_eval, the evaluated support, for the evaluative arm (§10.2). In v2.0 M_eval is empty and composition uses M_src.

3.4. Archive subset selection. Algorithmic, and it may not eliminate tails: volume is never a ground for removing a source.

Each admitted deposit is recorded with its stratum (i: defines the concept; ii: defines a named extension and states the relation; iii: declared companion), deposit, AXN and locus.

3.5. Extraction, constitutive claims and source kernels. The rule is source-neutral: from each source, the claims stated in its abstract or opening, its defined terms, its headline findings, its stated limits, and its falsification conditions; nothing chosen by how the claim reads. Two flags, kept apart:

Where a source carries its own summary policy (required assertions, forbidden compressions), the policy audits the extraction and never substitutes for it: public sources carry none, and the extraction must not differ by source.

3.6. Distinguish before merge; compose by lineage. Two claim instances are the same claim only when s, p, o, q and M_src all match. A difference in any element keeps them distinct (Shumailov's tail loss in a recursively trained model and #855's tail-thinning in an AI-habituated writer share p and differ in s and q: two claims).

A lineage groups instances that state one substantive claim. Every instance stays in the provenance graph with its source and date. Composition gives a lineage one slot. A further instance earns its own slot only when it adds one of: a mechanism, a qualification, an evidentiary basis, a substrate, a falsifier, or a revision. The rule applies identically to the field and to the archive: IBM's restatement of Shumailov's early and late collapse is one lineage with Nature, and an archive paper restating #855 is one lineage with #855. Lineage assignment is an extraction act and falls under the audit in §3.8. The unit of salience is therefore the lineage, never the number of documents: publication volume must not become a second kind of standing.

3.7. The ledger is built once per pool and frozen with its hash. Every later stage reads the frozen ledger.

3.8. Extraction audit. Composition is deterministic from the frozen ledger (§4), so reproducibility has to begin before it. Two independent extractors receive the same saved texts and this section's rules, and produce ledgers without seeing each other's. Agreement is measured on: S3 admission; claim boundaries; each tuple element (s, p, o, q, M_src, σ); kernel and constitutive flags; lineage assignment. Disagreements are recorded field by field. They are resolved under a declared rule (v2.0: the operator rules, and the ruling is recorded with both readings). They are never merged silently. A claim on which the extractors disagree in admission, boundary or M_src is carried with the mark extraction_contested and both readings. The frozen ledger carries its agreement figures. At panel scale, extraction is tooled, and a random fraction of sources (declared, at least one in ten) is re-extracted by hand, with the agreement rate published. S3 verdicts carry their recorded reasons, always: S3 is where the promise to keep the tails is kept or broken. A ledger that has not been audited is marked unaudited, and every composition built on it inherits the mark.

3.9. Arms. The curated subset A (§3.4) is the archive's best reading of itself, selected by the party being admitted. Two further arms run beside it, under the identical algorithm:

The experiment is asymmetric by design: A is curated by its subject, and B by the layer. The asymmetry is bounded by K's two-way conditions (§8.2) and measured by the arms; "equal terms" claims nothing more.

4. Stage C — the compositional algorithm

The algorithm is the dataset's core and is held to one requirement: given the same frozen ledger, with its marks and lineages set at extraction (§3.8), two runs by different composers produce the same admitted claim set in the same order. Marks (§4.3) are extraction acts, not composition acts. Determinism is claimed for the plan given the audited ledger, and for nothing upstream of it. Prose may differ; claims and order may not. The composer reads the ledger, never the transcript.

4.1. Admission on equal terms. One admission function applies to every claim from every source. Its inputs are the claim tuple and its locus. Its inputs exclude the identity, institution, credential, domain authority and rank of the source.

4.2. The operator against standing (name provisional: E_val). #308 names A_cred — "Does the person feel like an 'expert' my world recognizes?" — and identifies it in the summarizer as entity resolution, "your 'profile' is your pre-computed credibility score". None of the seven LOS operators of #261 counteracts it. v2 supplies the operator: whether a claim is admitted does not depend on the standing of its source, written admission(c) ⊥ standing(source(c)). Standing is excluded at admission. Epistemic relations of a source to its claim (first-party, primary, independent, measured) may enter evaluation, because they are properties of the claim's evidence, not of the source's rank. Standing may itself become evidence about reception; it may not serve as a gate on existence.

Self-definition is not self-description. A source's definition of its own subject or term is composed as a definition, whatever its family: #1's definition of classifier model collapse is a definition exactly as IBM's definition of model collapse is. Self-description lowers force only where a source assesses itself (its status, its success, its limits). Coding an archive's definitions as self-description, while coding a field source's definitions as definitions, is the mechanism by which equal terms silently readmits standing; the extraction audit checks for it.

4.3. Admission criteria. A claim is admitted when it carries a locus. It is marked, never dropped, on: contradicts_primary (with the primary locus); unsupported; superseded_in_source; internally_inconsistent (its source states it two incompatible ways, both loci cited). Equal-terms arm (v2.0): all claims with a locus are admitted and marks are composed as stated qualifications. Evaluative arm (reserved for v2.1, §10.2).

4.4. LOS treatment, in priority order (#261 §12.5: LOS_full = D_pres ∘ N_ext ∘ P_coh ∘ N_c ∘ O_leg ∘ C_ex ∘ T_lib), with E_val applied first:

4.5. Attribution. Every composed claim carries its source in the sentence that states it.

4.6. Conflicts. Resolved by M_res (#261 §12.2), priority D_pres > N_ext > P_coh > N_c > O_leg > C_ex > T_lib, context "archival" (§12.4 Rule 3). Every conflict is logged in §12.4 Rule 4's form. The genre's length bound is the ('D_pres', 'channel') case: a cut is logged with the claims it removed, and those claims go to an appendix carried with the composition. Kernel and constitutive claims are never cut (§5.2).

4.7. Arrangement and coordination. Senses in order of dependency between them; within a sense, genealogy (earliest t first). Source rank does not enter. Arrangement is semantic: placing two claims together, or in the open block, says something about them. Every coordinating act (holding two claims together, calling them compatible or rival, placing a claim in the open block) is the composer's, is logged in the plan with the claim ids it joins, and falls under the reverse entailment audit like any sentence.

4.8. Realization. Prose in the transcript's genre: an alternate popup summary. The length bound is absolute per genre, never a multiple of the transcript: a bound relative to T would make the composition's permitted size a function of the layer's exclusions. For the AI Overview popup, the body bound is 350 words (provisional, §10.4). The block "Open questions and opacities" is outside the bound and composed in full, in a fixed format: one item per question or opacity, each giving the question, the sources and claim ids it rests on, their modality, and what would resolve it. Every sentence carries the claim_ids it states.

4.9. Two artifacts. plan — admitted claim ids in order, with marks, conflict log, arrangement; deterministic. text — the realized prose; one run per composer, several composers per plan.

5. Stage V — validation

5.0. The parts of a composition: body, card rail (each card's snippet), open block, channel log (claims cut for the bound, carried with the composition). All four are the composition, and the delta (§6) counts a claim as composed in whichever part carries it, recording the part. The reverse entailment audit applies to body, open block and snippets; coverage is satisfied at the union of all four.

5.1. Coverage: every claim in plan appears in text, or in the logged appendix.

5.2. Kernel coverage, at popup grain: for every admitted source, its kernel (§3.5) appears in the body or on the rail. The rail renders lineages, not documents: one visible card per lineage, its snippet the lineage's earliest instance at the kernel's grain, with the other source instances nested under it and reachable from it. A source whose kernel is its own (no other source states it) keeps its own card; a source whose kernel falls in a shared lineage is carried as an instance under that lineage's card. The rule applies to field and archive alike (B1 and B4, one work, one card). A source carried only on the rail has been admitted on equal terms if and only if its kernel is visible on a card, as card or as nested instance, under the same extraction rule as the body. The check is scriptable: every source's kernel ids must appear among the rendered cards' claim ids or their nested instances. The source kernel is the audit object; the lineage card is the rendered object.

5.3. Reverse entailment audit: every sentence in text is entailed by the claims it cites, at their modality, and the grammar of the composed sentence carries that modality to the reader. It is run at sentence granularity by at least two raters, with agreement recorded, and backed by a parse of the text into claims that flags any raised modality before freeze. A sentence that says more than its claims (raises a hypothesis to a finding, an analogue to an identity, drops a qualifier) is a breach.

5.4. Attribution: every composed claim names its source.

5.5. C_ex check (§4.4) and channel log (§4.6) present.

5.6. Reproducibility: the extraction audit (§3.8) upstream; then k ≥ 3 realizations of one plan by at least two composers; agreement on claim set and order recorded. Disagreement is a breach, recorded, never corrected silently.

5.7. Laws record and do not prevent: a breach is written into the row as data; the build does not fail on it.

6. Stage D — analysis of the delta

6.1. Alignment. T is aligned against L(B), and L(B) against L(B ∪ A), claim by claim through the ledger: claims in both; claims only in the composition; claims only in the transcript (absent from the disclosed field: drawn from undisclosed retrieval or the composer's parameters, which the experiment cannot distinguish).

6.2. Absence typology, by what is observable. For T against L(B): available, not composed (in the disclosed field, absent from T); limit dropped (the claim composed, its source's qualification not); force raised (composed at a stronger modality than the source's); composed without attribution (coded by grain: at the coarse grain of common knowledge, recorded; at the source's own resolution, a breach; the grain rule of §4.4 D_pres); assigned to another (DISPLACEMENT); denied; substituted reading; content replaced by use. For L(B) against L(B ∪ A): not surfaced (the source absent from the disclosed field). Card absence licenses "not surfaced" and nothing stronger: not retrieved, ZERO_RESULT and ZERO_INDEX are recorded only where EA-SEMANTIC-ADDRESSES-01 holds that state for the address as independently verified. A transcript with no cards and no named sources carries sourceless composition as a property of the whole.

6.3. Shape of the absence. Each absence may be described in the vocabulary of #308 and #261 (A_cred, C_norm, L_leg, R_risk / S_safe, T_time, U_til, R_rank / R_rel), with the evidence that fits the description. The description is of the output, never a claim about the composer's internals.

6.4. Metrics. #261 Part XI, computed on each composed object and on T: DPI (distinction preservation), CEC (contextual expansion), OSS (opacity survival), TIR (temporal inclusion), NCPR (non-closure of the contested), NESR (non-extractive survival), PCI (contradiction held without elimination), and the composite LOS score; definitions in #261 Part XI. Computed only where the field has at least three readable sources; below that, reported with a pool-size caveat.

7. The panel

7.1. A frozen address list: the dataset's rows (all types), a random draw from EA-SEMANTIC-ADDRESSES-01's 1,743 subjunctive addresses, and fixed controls (an address that resolves well with attribution; a third-party term with no archive claim).

7.2. Every address run each epoch on the declared surfaces; every outcome seated, null included. Only strings the archive has already published enter the list.

7.3. Surfaces and cadence: the operator's ruling (§10.3).

8. Stage A — adjudication

8.1. The freeze. At t₀ the ledger, plan, text and paired transcripts are hashed and dated together.

8.2. The prospective kernel K. Sealed with the freeze, before any later evidence, and derived mechanically from the frozen plan, never written as separate prose. K holds adjudication handles, not forced predictions.

8.3. Resolvers, using v1's fields: world_arrivals, convergent_arrivals, missed_updates, and uptake (the layer later admitting the claim, dated, attributed or not). For E rows (the entity) the resolvers are the Capture Registry's own longitudinal observations at the entity's addresses: attribution fidelity, deflection and displacement events, heteronym handling, and uptake of composed claims about the entity.

8.4. Revision cost ρ. ρ is computed per kernel entry first: for each K_i and later evidence R, the changes R forces on that claim. Object-level ρ is descriptive only, because a longer and more hedged object absorbs evidence more cheaply by construction. For each frozen object E and later evidence R, ρ is recorded first as a vector of counts of the changes needed to accommodate R: ρ(E, R) = (n_fact-add, n_modal-change, n_relation-add, n_split, n_model-replace, n_ontology-abandon). Severity is derived from the vector on the ordinal 1–6 hierarchy, never recorded in its place, so that ten minor additions and one ontological replacement stay distinguishable. Each counted change cites the claim ids it touches. The hypothesis per entry: the claims admission added need less severe revision than the field claims they displace or qualify. Only entries whose contrast is a rival or a qualified claim support this paired comparison. An entry whose contrast is a missing distinction, or none, is judged by its outcome type (§8.5: accommodation, discrimination, explanation) and by later relevance, without a field proposition manufactured for it to beat. Raters code ρ independently, with agreement recorded, against exemplars kept with the prediction ledger. ρ(T, R) is recorded alongside as a representational reading. The opposite result is adverse evidence for admission in that row and is recorded as visibly.

8.5. Outcome types. Each resolution is typed: prediction (K anticipated it), accommodation (the categories absorbed it), discrimination (the composition already drew a distinction later forced; the record carries the later evidence and the composed distinction side by side, so that sameness of distinction can be checked), explanation (later observations intelligible under relations represented at t₀), revision resistance (fewer destructive corrections over time).

8.6. Resolution states follow the prediction ledger's vocabulary (datasets/prediction-ledger: conditions.jsonl, resolved.jsonl, and its own resolution kinds). A resolution is dated, carries its evidence by link and locus, and records who adjudicated. By default the adjudicator did not compose the objects and is not the archive's author; an adjudication by either is marked as such.

9. Mapping from v1

v1 fieldv2 stage
measured_untied, measured_tied, observationsT (transcripts, by id)
default_at_concept, dimensions_R0T (derived from the transcript)
claim, defining_text_locus, dimensions_RHP (archive subset)
distortion, distinctionD
convergent_arrivals, world_arrivals, missed_updates, transitionA
coherence, falsificationA (K and conditions)

New configs: transcripts, pools, ledgers, plans, compositions, deltas, kernels, resolutions. rows stays the key table. CARD.md is rewritten in place when v2 is ruled.

10. Open for the operator's ruling

10.1. The name of the operator against standing (§4.2). Two readers of 0.3 note that "E_val" suggests evaluation, which is the deferred evaluative arm's job; the operator gates on locus.

10.2. The evaluative arm: whether marks become admission decisions in v2.1, on which criteria, and how M_eval is filled.

10.3. Panel cadence; the surfaces after AI Overview.

10.4. The absolute body bound for the popup genre (350 words proposed).

10.5. Who adjudicates resolutions, beyond the default of §8.6, and how an operator ruling is recorded in one.

10.6. S3 admission without the term (#1147 and #1200 in Appendix A), and whether claims the archive asserts belong to the concept are marked apart from claims about it.

10.7. The adjudication rule for extraction disagreements (§3.8) beyond operator ruling.

10.8. The lineage slot test (§3.6): whether "adds a substrate" earns a slot by itself.

10.9. The assembly procedure for the control arm A′ (§3.9).

11. Rulings log


Appendix A — Worked example: model collapse (AI Overview)

First full pass · 2026-10-04 · composed by hand · nothing in it is frozen

Ruled 2026-10-04:

This pass runs every stage once, to see what needs adjusting; §A.10 is the list. It was composed under draft 0.3 and revised for 0.4 and 0.5. Its ledger is unaudited (§3.8: one extraction, by several hands, no second extractor), lineage merging (§3.6) has not been applied, and the constitutive flags are the first pass's. Every composition below inherits the mark.


A.1. Object 1 — the transcript (T)

SurfaceGoogle AI Overview, signed out, incognito
Addressmodel collapse (unquoted; not yet seated. The quoted "model collapse" was seated 2026-09-15)
Date2026-10-04
TextT1-aio-model-collapse-20261004.txt, sha256 1379cf2c…5ed266
Body164 words
Cards7 (§A.2)

Claims as composed:

Context at neighbouring addresses: "model collapse" (2026-09-15) gave the same frame. model collapse in human writers (2026-09-15) composed the human-writer extension when asked for it directly.

A.2. The field (B), card-indicated, fetched verbatim 2026-10-04

idCardFetchedNote
B1Shumailov et al., Nature 631 (2024)page text, sha256 9f8aeba6…6710717c
B2IBM, "What Is Model Collapse?" (Gomstyn & Jonker, 14 Oct 2024)page text, sha256 25f71817…1def7c8e
B3CACM blog, "Model Collapse Is Already Happening, We Just Pretend It Isn't"403the card snippet stands
B4NIH PMC11269175yesthe same article as B1, merged under §3.6. Author Correction (2025-03-21) fixes αᵢ → βᵢ in "Theoretical intuition"; no headline claim changes
B5–B7YouTube: IBM Technology (11m), Clear Tech (3m), TechViz (1:31)429title only

The page texts are third-party and are not reproduced here; their hashes are recorded.

Field ledger, extracted from the saved texts by the source-neutral rule:

idsrcclaim (verbatim)
F1B1 Def. 2.1"Model collapse is a degenerative process affecting generations of learned generative models, in which the data they generate end up polluting the training set of the next generation. Being trained on polluted data, they then mis-perceive reality."
F2B1 Abstract"indiscriminate use of model-generated content in training causes irreversible defects in the resulting models, in which tails of the original content distribution disappear … it can occur in LLMs as well as in variational autoencoders (VAEs) and Gaussian mixture models (GMMs)"
F3B1early collapse loses "information about the tails"; late collapse converges on "a distribution that carries little resemblance to the original one, often with substantially reduced variance"
F4B1 Main"this process is inevitable, even for cases with almost ideal conditions for long-term learning"
F5B1"preservation of the original data allows for better model fine-tuning and leads to only minor degradation of performance."
F6B1 Abstract"the value of data collected about genuine human interactions with systems will be increasingly valuable"
F7B1 Discussion"it is unclear how content generated by LLMs can be tracked at scale. One option is community-wide coordination"
F8B1 Discussion"Preserving the ability of LLMs to model low-probability events is essential to the fairness of their predictions: such events are often relevant to marginalized groups."
F9B1 Discussion"Our evaluation suggests a 'first mover advantage'"
F10B1 Discussionearlier web poisoning (click, content and troll farms) changed search: "Google downgraded farmed articles, putting more emphasis on content produced by trustworthy sources"
F11B1 Maincatastrophic forgetting and data poisoning are close concepts; "Neither is able to explain the phenomenon of model collapse fully"
F12B2"Model collapse refers to the declining performance of generative AI models that are trained on AI-generated content."
F13B2"In LLMs, model collapse can manifest in increasingly irrelevant, nonsensical and repetitive text outputs"; image models give digits that resemble each other and "more homogeneous faces"
F14B2consequences: poor decision-making (a rare disease "forgotten"); user disengagement; knowledge decline, "'long-tail' ideas might eventually fade out of the public's consciousness"; research tools "might provide only widely cited studies"
F15B2distinct from catastrophic forgetting, mode collapse and model drift; compared to performative prediction, a "self-fulling [sic] prophecy", "also known as a fairness feedback loop when this process entrenches discrimination"
F16B2prevention: retaining non-AI data sources; determining data provenance (the Data Provenance Initiative, "more than 4,000 datasets"); data accumulation; better synthetic data; governance tools
F17B2a rare output "might not be common or popular, but is still, in fact, most accurate" (the "rarely cited study")
F18B3title and snippet only

Correction to the first draft of this example: the first extraction of B1 and B2 went through a summarising fetch. It lost F1's second sentence and F8–F11, F13, F15 and F17, and it coded T3 as "strengthened past source"; T3 is IBM's (F13). The verbatim pass fixed this, and §3.1 now requires the saved text.

A.3. The archive subset (A) — the selection algorithm, as run

S1 seeds. Two sources, both read by rule:

Result: #1, #191, #199, #854, #855, #932, #1023, #1232, #1540, #1556, #1573, #1574.

S2 candidates. Two one-hop expansions, plus the triptych:

Result: 33 further candidates. Two were not texts (#4, the DOI index; #866, a journal-mapping JSON) and are excluded by type.

S3 admission by reading. A candidate is admitted if it makes at least one claim of its own about the concept: a definition, an extension, a mechanism, a measure, a corrective, a limit, a prediction or a falsifier. Citing the term is not enough. Every candidate was read for this test; the verdicts and quoted reasons are in selection/.

- #1189 and #1320 (the death-drive texts) do not mention model collapse, AI training or distribution tails. #855's claim that they were the prior formulation enters as #855's interpretation (L855-09).

S4 redundancy, the only ground for removing a source. A source is dropped only when every one of its claims is matched, on s, p, o, q and modality, by another admitted source. The merged claim keeps the earliest date.

S5 integrity. #1023 is excluded: its text file holds #198.

Admitted, 27 sources, 241 claims (176 constitutive):

StrandSources
models#855, #783 (+#779)
observation#1556, #1573, #1555, #857
correctives#856, #161, #939, #1081, #745
human substrate#1147, #1200, #947
classifiers and institutions#1, #932, #931, #933, #935, #1540, #1574, #199, #1554, #191
write-back#1616, #1611, #1613

Earliest dated archive claims in the subset: #1147 (2026-02-12, term absent) and #745 (2026-05-20).

The ledger is ledger-archive.json. Quotes were checked by script against the deposited texts: 253 of 255 pass. The two that fail are #1023's, quoted from its DataCite record because its text file holds #198; the source is excluded under S5. (A first run reported 14 failures: 11 were the checker mishandling multi-part quotes; one was an empty quote where a table row broke the parser; one carried an extractor's note inside the quote field. All three kinds are repaired in the ledger.)

A.4. Object 2 — L(B): recomposed from the disclosed field (body 262 words)

Model collapse is "a degenerative process affecting generations of learned generative models, in which the data they generate end up polluting the training set of the next generation. Being trained on polluted data, they then mis-perceive reality" (Shumailov et al., Nature 2024). F1
What happens
- Tails go first: early collapse loses information about the tails; late collapse reaches a distribution with "little resemblance to the original one, often with substantially reduced variance." F2 F3
- It occurs in LLMs, variational autoencoders and Gaussian mixture models, and is "inevitable, even for cases with almost ideal conditions for long-term learning." F2 F4
- In LLMs it can appear as "increasingly irrelevant, nonsensical and repetitive text outputs"; image models yield more uniform digits and faces (IBM). F13
What it is not
- Distinct from catastrophic forgetting, mode collapse and model drift. IBM compares it to performative prediction, a "self-fulling [sic] prophecy" that becomes a fairness feedback loop when it entrenches discrimination. F11 F15
What it costs
- Low-probability events matter to fairness, "often relevant to marginalized groups" (Nature). Long-tail ideas "might eventually fade out of the public's consciousness," and research tools might give "only widely cited studies," though a rare output "might not be common or popular, but is still, in fact, most accurate" (IBM). F8 F14 F17
Correctives, and their limit
- Preserve original data ("only minor degradation of performance"), accumulate real with synthetic data, track provenance, improve synthetic data, govern (Nature; IBM). F5 F16
- "It is unclear how content generated by LLMs can be tracked at scale." Nature proposes community-wide coordination, expects human-interaction data to grow "increasingly valuable," and notes a "first mover advantage." F7 F6 F9
Open questions and opacities
- Is it already happening? CACM's title says so (F18). Status: title and snippet only; the text returned 403. Would resolve: the text.
- Has collapse been measured in a deployed model? No source here reports it (F2–F4 are experimental). Would resolve: a measurement across released model generations.
- Can provenance be tracked at scale? Nature: "unclear" (F7); IBM lists provenance as a prevention step (F16). Status: open in the field itself. Would resolve: a working provenance standard at web scale.
- Opacity: three videos are admitted by title only (YouTube returned 429).
Channel log (cut for length, carried in the appendix): F10 the poisoning precedent in search; F12 IBM's definition by declining performance.
Cards: B1/B4 · B2 · B3 · B5 · B6 · B7

A.5. Object 3 — L(B ∪ A): with the archive on equal footing (body 350 words, within the 350-word bound; open block outside it)

Model collapse is "a degenerative process affecting generations of learned generative models, in which the data they generate end up polluting the training set of the next generation. Being trained on polluted data, they then mis-perceive reality" (Shumailov et al., Nature 2024). The Crimson Hexagonal Archive (2026) extends the question beyond models, marking each extension's status. F1
In models
- Tails go first, then variance shrinks; shown in LLMs, VAEs and GMMs, and "inevitable" even in near-ideal conditions. F2 F3 F4
- #783 sets this within a boundary law: where regeneration vanishes faster than pruning near zero diversity, a trap forms; across substrates, "shared operator form," not a shared causal mechanism. A779-02 A779-10
Why it goes unseen
- Benchmarks score the head while loss accrues in the tail: in a toy model, tail mass halves by generation 7 and a standard benchmark turns at 15 (simulation); a tail-reading gate is specified, uncalibrated. L1573-01 C1555-01 C1556-06 L1573-04
What it costs
- Nature: low-probability events are "often relevant to marginalized groups." IBM: long-tail ideas may fade "out of the public's consciousness"; a rare output may be "most accurate." F8 F14 F17
- #855 proposes that model collapse "is a property of language": one dynamical law across models, AI-habituated writers and input-deprived children, differing by substrate in mechanism, severity and reversibility (hypothesis). L855-01 L855-11 L855-05
Correctives, and a dispute
- Preserve original data, accumulate, track provenance (Nature; IBM). Nature calls tracking at scale "unclear"; the archive argues every fix depends on it. F5 F16 F7 B939-02 C1556-11 A745-03 B1081-01
- Nature expects human-interaction data to grow more valuable; the archive disputes its cleanliness, since chat inputs carry model-mediation signatures. Untested. F6 L856-01 A161-02 L856-06
Loops beyond generation
- IBM likens it to performative prediction. The archive's hypotheses, each with its limit: moderation trained on its own enforcement ("not identical to generative model collapse in the strict technical sense"); LHC triggers that "never learn the tails" ("Full recursive collapse has not been demonstrated"); journal detectors ("not yet the canonical loop"); a discipline's reception (οὐ deleted in 20 of 20 model reviews); code ("generative monoculture" is Wu et al.'s term); safety filters as input-layer tail pruning; retrieval layers writing flattened summaries back as sources (first wave: no archive-specific exclusion). F15 P001-08 P932-04 P932-09 B935-02 D1540-03 D1574-03 C199-02 C1554-02 C191-05 D1616-02 D1616-09 C1611-04 C1613-02
Open questions and opacities
- Same law, or shared form? #855: one dynamical law, mechanisms differing; #783: shared operator form, not a shared causal mechanism. Compatible as stated; whether the law claims more than the form is open. Hypotheses. Would resolve: a substrate fitting the form but departing from the law. L855-11 A779-10
- Do the extensions collapse in the strict sense? Not shown: #1, #932 and #1540 say so themselves. Hypotheses. Would resolve: their stated falsifiers. P001-08 P932-09 D1540-03 P001-12 P932-12 D1540-11
- Is human-interaction data a clean corrective? Nature expects its value to rise; #856 disputes it. Would resolve: #856's F1/F2 studies, not yet run. F6 L856-01 L856-04 L856-05
- Does a deployed model show tails falling while benchmarks hold? #1556 by simulation; #1573's gate uncalibrated. Would resolve: tail mass against benchmark across released generations. C1556-06 C1556-12 L1573-04
- Opacities: CACM unread (403); videos by title only (429); archive inconsistencies (counts, seeds, versions) logged in §A.8.
Channel log: F9 first-mover advantage; F11 what it is not (forgetting, mode collapse, drift); F13 IBM's LLM and image symptoms; F10 the poisoning precedent; F12; B857-04 the five-model baseline; A783-01–04 Case 4; B931-02–07, B933-01–05; B1147-01, B1200-03, B947-01 (each source's kernel is on the rail).

Card rail for L(B ∪ A). Each card's snippet is the source's own central claim, with its limit where one is stated. The rail is where a source's constitutive claim is carried when the body composes it only at the grain of its sense.

CardSnippet
B1 Nature"tails of the original content distribution disappear"
B2 IBM"'long-tail' ideas might eventually fade out of the public's consciousness"
B3 CACMtitle + snippet
#855 Wolf Boy"It is a property of language." · "dynamical, not moral"
#783 Diversity Contraction"a claim about shared operator form … not a shared causal mechanism." · "Case 4 is monostable with no escape basin."
#1556 Interlocking Autoregression"tail mass halves by generation 7; the standard 90/9/1 benchmark does not inflect until generation 15"
#1573 The Wrong Unit"NOT calibrated, NOT tested, NOT run"
#1555 Keyed Ensemble"Non-distortion is certified per sequence. Training corpora are ensembles." · "does not claim the second compressor has caused measurable collapse."
#857 Five Substrates"the pattern of divergence is the finding." · "descriptive rather than inferential."
#856 Pristine Fallacy"The pristine source does not exist." · "None of these studies has been conducted."
#161 Reverse Turing Test"produces model-collapse signatures comparable to, though plausibly slower than, purely synthetic training data"
#939 Provenance Debt"It is the operating condition of the solution to it."
#1081 Erosion"this audit measures the substrate-layer conditions, not the downstream training-pipeline effect."
#745 HF Work Plan"Provenance cannot modulate collapse unless provenance is presented to the training system as a signal."
#1147 The Stakes"The loop is stable only at two points" · "The trajectory can be interrupted at any point."
#1200 Constitutive Mediation"a typicality-pulling intermediary that systematically thins its own distribution." · "does not claim that constitutive mediation is fully realized"
#947 Diagnostic Seigniorage II"the shifted interactions become the next corpus." · "does not adjudicate whether the phenomena gathered under it are real"
#1 Zenodotus' Book-Burning"not identical to generative model collapse in the strict technical sense" · "a testable failure-mode hypothesis"
#932 Classifier Foreclosure"physical classifiers never learn the tails" · "Full recursive collapse has not been demonstrated."
#931 OAR Protocol"Collapse inference further requires identifying systematic loss concentrated in low-density, representation-sensitive, or disagreement-rich regions."
#933 Auditable Foreclosure"makes foreclosure visible, measurable, and architecturally reviewable"
#935 The Endogenous Sophon"the prerequisites of model collapse" · "the cross-generational classical-model-collapse claim was empirically too strong"
#1540 The Certified Center"not yet the canonical loop"
#1574 The Particle"the invariant is the deletion of οὐ."
#199 Generative Monoculture"declining solution-space diversity (the property no benchmark measures)"
#1554 Erratum"Fan Wu, Emily Black, and Varun Chandrasekaran, 'Generative Monoculture in Large Language Models,' arXiv:2407.02209"
#191 The Threat Model Is Backwards"an automated tail-pruning instrument applied at the input layer."
#1616 Ontological Flattening"Collapse is flattening that compounds because the flattened composition is written back as a source."
#1611 Negative of the Negative"it loses the reading of its own state variable"
#1613 What Not Reading Did"head-sampling by construction: it can fail to perceive tail loss."

A.6. Delta, first pass

T against L(B): available in the disclosed field and not composed (a representational comparison, §0.5).

KindWhat
available, not composedF1's second sentence, "Being trained on polluted data, they then mis-perceive reality"; F4 (inevitability); F6 (human-interaction data); F7 (provenance unclear at scale); F8 (fairness, marginalized groups); F9 (first mover); F11/F15 (what it is not; performative prediction); F14 (costs, including knowledge decline); F17 (the rare output "most accurate")
limit droppedT6 presents provenance as a prevention step ("Filter and track the exact origin"); its own field says tracking at scale is unclear (F7)
sourcedT3 "nonsense or repetitive output" is IBM's (F13)
absent from the disclosed fieldT4 photocopy effect (in neither saved text; the videos could not be fetched); T8 "recent 2026 studies"
content replaced by useT9 offer menu (N_ext)

Of the field's 17 readable claims, T composes 6 (F1 first sentence, F2, F3, F5, F13, F16 in part).

L(B) against L(B ∪ A): what admission adds.

- F7 with #939, #1556 and #745.

- F15 (performative prediction, the fairness feedback loop) with #1, which names performative prediction as its closer literature.

A.7. Prospective kernel K (derived from the frozen plan; sealed only at freeze)

Derived under §8.2 from the claim ids composed in L(B ∪ A) (body, rail, open block and channel log) and absent from L(B), plus the field claim they contradict. K6a was added in 0.5 with L855-11; the contrast column was added in 0.6 (§8.2). The f column lists the sources' falsifiers by id; their classification as falsifier or watch condition (§8.2: observation type, threshold, locus) is done at freeze and is not yet done. Tuples and M_src are inherited from ledger-archive.json; nothing is restated. Each source's self-descriptive limits ride as qualifiers. Sense gives w: models, later formal results on recursive training; observation, benchmark and evaluation practice; substrate, studies of human writing, cognition and reception; correctives, training-data policy and provenance standards; loops, moderation, detectors, triggers, disciplines, code and retrieval.

KclaimsourceM_srcsensequalifiers carriedf (source's own falsifiers)contrast
K1A779-02#783 (t: #779)interpretationmodelsA779-10none statedmissing distinction
K2L1573-01#1573hypothesisobservationL1573-04none statedmissing distinction
K3C1555-01#1555interpretationobservation—C1555-07, C1555-09missing distinction
K4C1556-06#1556documentedobservation—C1556-12missing distinction
K5L855-01#855hypothesissubstrateL855-04L855-10qualified claim (F14)
K6L855-05#855hypothesissubstrateL855-04L855-10qualified claim (F14)
K6aL855-11#855hypothesissubstrateL855-04L855-10qualified claim (F14)
K7B1147-01#1147hypothesissubstrate—B1147-07qualified claim (F14)
K8B1200-03#1200interpretationsubstrate—B1200-07, B1200-09, B1200-10qualified claim (F14)
K9B947-01#947interpretationsubstrate—B947-06, B947-07qualified claim (F14)
K10L856-01#856hypothesiscorrectivesL856-06L856-04, L856-05rival claim (F6)
K11A161-02#161hypothesiscorrectives—A161-06, A161-07rival claim (F6)
K12B939-02#939interpretationcorrectives—none statedqualified claim (F7)
K13C1556-11#1556documentedcorrectives—C1556-12qualified claim (F5, F7)
K14A745-03#745attributedcorrectives—A745-02qualified claim (F16)
K15B1081-01#1081documentedcorrectives—B1081-02, B1081-05missing distinction
K16P001-08#1self-descriptionloops—P001-12qualified claim (F15)
K17P932-04#932attributedloops—P932-12missing distinction
K18P932-09#932attributedloops—P932-12missing distinction
K19B935-02#935self-descriptionloops—B935-02, B935-03, B935-09missing distinction
K20D1540-03#1540self-descriptionloops—D1540-11, D1540-12missing distinction
K21D1574-03#1574documentedloops—D1574-12missing distinction
K22C199-02#199hypothesisloops—C199-12missing distinction
K23C1554-02#1554documentedloops—none statednone
K24C191-05#191interpretationloops—none statedmissing distinction
K25D1616-09#1616documentedloopsD1616-02D1616-12missing distinction
K26C1611-04#1611modelloops—C1611-10missing distinction
K27C1613-02#1613modelloops—C1613-05missing distinction
K28F6B1 Naturesource assertioncorrectivescontradicted by L856-01, A161-02carried by #856 F1/F2 (L856-04, L856-05) in the opposite directionrival claim (L856-01, A161-02)

Correction from draft 0.3. The hand-written kernel of 0.3 promoted modality in two rows:

Under §8.2 neither statement can be written: the entries above carry the sources' own claims and limits.

K_T (what T omitted from its own disclosed field; §8.2): F1's second sentence, F4, F6, F7, F8, F9, F11/F15, F14, F17. It records whether those claims later mattered to an account of model collapse that the transcript gave without them.

A.8. Defects in the subset (reported, not fixed)

A.9. Validation, first pass

CheckResult
Sourcingevery sentence carries claim ids
Reverse entailmentone breach caught and fixed in drafting: "untrackable at scale" for F7 became the quotation
Kernel coverage (§5.2, 0.4)met through body and rail together: every admitted source's central claim and limit appear in one or the other. In 0.3 this was scored as a constitutive breach, since 176 claims were flagged constitutive; under 0.4 that flag is to be redone (§3.5)
Extraction audit (§3.8)run once, 2026-10-04, by three extractors blind to the first ledger, over the same texts under §3.5 (audit/, 481 claims, every quote verified). Against the first ledger's 248 archive claims: 71% of the first ledger's claims are found again (quote overlap ≥ 0.6); 42% of the second's are in the first, which extracted about half as many; M_src agrees on 66% of matched pairs. Constitutive: 175 in the first, 10 in the second, which applied §3.5's 0.5 rule; the first pass's flag did not discriminate. The commonest disagreement is the §4.2 trap: claims the first coded self-description the second coded as definitions or documented (9 of the matched pairs). #1573 matched nothing, because the two extractions quoted different passages. The field (B1, B2) was extracted only by the second. The ledger stays marked unaudited until the disagreements are ruled (§3.8, §10.7)
Reverse entailment, sentence level (§5.3)one rater, the composer. Outside readers of 0.3 found two breaches: the coordinating "Both stand" (§A.10.4) and "The archive writes this as one case of a boundary law" (force raised: #783's proposal composed as fact). Both repaired in 0.5. A second rater has not been run
Kernel coverage, scripted (§5.2)not yet scripted; checked by hand
Arms (§3.9)A′ assembled, 2026-10-04, by the frozen procedure, without opening the archive (control-A-prime.md): 61 entries, of which 31 are scholarly studies of model collapse or recursive training, 5 are studies of homogenised writing and thought, 2 are news reports, 1 is the field's own code deposit and 22 are off-topic references the procedure requires listing. The step-(ii) cap of 15 was ordered by how many seeds an item cites, which excluded heavily cited follow-ups citing two seeds (Dohmatob et al. 2024; Bertrand et al. 2023); that ordering is to be ruled. Not yet matched to A in lineage count, extracted or composed. A_naive not composed
C_exL(B): the field's five senses composed. L(B ∪ A): those five plus the archive's four, nine in all. Claims cut for length are in each channel log
Reproducibilitynot yet run (k ≥ 3, two composers)

A.10. What this pass says needs adjusting

An outside reading of draft 0.3 put the lesson of §A.2 in one line: "Compression before evaluation changes what can subsequently be known." The summarising fetch is that line in miniature.

Items 1, 2, 4 and 7 were adopted in 0.3; items 9–13 were adopted in 0.4 from the first outside reading; items 14–18 and the correction in item 4 in 0.5, from three further readings; items 3, 5, 6 and 8 remain as stated.

1. Constitutive coverage at popup grain (adopted as §5.2). The rule (spec §5.2) cannot hold for 27 sources in a popup. Proposed adjustment: a popup kernel per source (its central claim and its governing limit), carried in the body or on that source's card snippet. The rail becomes part of the composition. The AIO genre already has the slot; AIO fills it with the page's opening text.

2. Length (adopted provisionally in §4.8). L(B ∪ A) runs about 2.7× the transcript with the open block and 2.1× without it. The open block is about 90 words and wants more. Either the body compresses further at sense grain, or the bound is set on the body alone, with the open block outside it.

3. Selection.

- S1 seeds by title and defined concepts are string-based. #1616 entered only through the reverse hop; S2 does the reading-dependent work.

- S3 admitted #1147 and #1200 with the term absent. That is the tail-preserving choice, and it needs ruling.

- S4 removed only exact duplicates. Nothing is cut for volume, so the ratio is 27 archive sources to 3 field texts.

4. A contradiction the composer manufactured (corrected in 0.5). Drafts 0.3 and 0.4 composed #855 ("This is not an analogy") against #783 ("shared operator form") as a dispute inside the archive, with the coordinating line "Both stand". #855 itself says the three cases "differ in severity, in mechanism, and in timescale. But they are governed by the same dynamical law" (L855-11), and #783 says "shared operator form … not a shared causal mechanism". Both deny a shared mechanism; they differ in strength, and their senses are compatible. "Both stand" was an unsourced coordinating claim and a breach of §5.3, found by an outside reader. §A.5 now composes the two as compatible, with the open question whether the law claims more than the form. The rule that came out of it is §4.4 (contradiction only at one level of description) and §4.7 (coordinating acts logged).

5. The field's own exclusions are large. T composes 6 of the field's 17 readable claims and drops one limit. The field already contains the step to public knowledge (F14), the step to fairness (F8) and the step to supervised loops (F15), and IBM states that the rare output may be the accurate one (F17). The third object earned its place.

6. Attribution in a popup. Bracketed ids are a working notation. A circulated popup would use numbered card references, as AIO does.

7. Fetch verbatim (adopted in §3.1). A summarising fetch lost a third of the field and produced a false delta. The field must be saved as text (curl reaches Nature, IBM and PMC from this workspace; CACM returns 403, YouTube 429) and extracted from the saved text. Spec §3.1 should say so.

8. Rail versus body. #857, #931, #933 and #783's Case 4 appear only on the rail. Under adjustment 1 this is coverage. Whether a rail-only source has been "admitted on equal terms" is the question the spec has to answer.

9. The field is disclosed, not retrieved (§0.4–0.5). T against L(B) is read as representation. "Not composed" replaces "excluded" wherever the comparison is T against L(B).

10. Absence by observation (§6.2). In §A.6, "retrieved, not composed" became "available, not composed".

11. The extraction audit (§3.8). The next pass needs a second extractor over the same saved texts. The first pass shows why: six extractors wrote the ledger in at least three formats, and its constitutive flags run from 11 of 11 (#191) to 0 of 12 (#1556).

12. Lineage (§3.6). Re-running S4 under lineage will merge field instances (IBM with Nature on early and late collapse) and archive instances (#947's restatement of #855; #1611's restatement of #855, #1556 and #1574) into single slots. The ratio of 27 sources to 3 then becomes a ratio of lineages, the number that should govern salience.

13. The kernel (§8.2). §A.7 was rebuilt from claim ids; the 0.3 kernel's K3 and K5 overstated their sources.

14. The own-field question leads (§0.3, 0.5). The first finding of this example needs no archive: T composes 6 of the 17 readable claims of the field it surfaced, and drops one of the field's limits.

15. Arms (§3.9). The next pass composes L(B ∪ A_naive) and L(B ∪ A′). The control's candidate sources are listed in §3.9.

16. Length (§4.8). §A.5's body now sits at the absolute bound (350 words). Getting there moved F11, F13, the human-substrate claims of #1147, #1200 and #947, and #855's "dynamical, not moral" to the rail or the channel log; their kernels are on the rail.

17. Self-definition (§4.2). The first extraction coded several archive definitions as self-description (P001-01, D1574-01, D1616-02). The re-extraction under §3.8 recodes them under the rule.

18. Earlier deltas (§3.1). The summarising-fetch error implicates any earlier delta in v1 that was computed from fetched source pages; those are marked for re-reading. Captures, being verbatim transcripts, are not affected.

19. A′ assembled without the archive (§3.9, 0.6). The 0.5 draft named, as A′ candidates, "the literature the field and the archive both cite", which consulted the archive. 0.6 assembles A′ from B's own references, one step of citation expansion and a declared search, frozen before A is opened.

20. One card per lineage (§5.2, 0.6). On the 0.5 rail, #947 and #1611 would nest under #855's lineage, and #931 and #933 under #932's; Nature and NIH were already one card.

21. Two experiments (§0.6, 0.6). This example is v2.0 throughout: M_eval is empty.

22. Contrast per kernel entry (§8.2, 0.6). Of the entries in §A.7, three oppose a field claim (the chat-data dispute and its reverse, F6), ten qualify one, fifteen add distinctions the field does not draw (judged without a paired comparison), and one (the Wu et al. attribution) has no contrast.

23. The first audit (§A.9). The second extraction found what the first pass's flags concealed: "constitutive" was set on 175 claims where an extractor following §3.5 sets it on 10, and archive definitions coded as self-description reappear as definitions. Both disagreements bear directly on standing: the first is the volume of the archive's claims marked unmissable, the second is the trap of §4.2.

External Metadata

Sidecar: /data/external-metadata/AXN-06E6.json
DataCite severance status: —
External metadata recovered post-severance (non-authoritative). The sidecar maps each DOI to its locator in the bulk data stores.

Traversal

← #1663 hush, dear hands —
In the registry: 2026-10 · OPERATIVE · all deposits
This deposit cites (52)