{
 "row": "model-collapse",
 "address": "model collapse",
 "epoch": "2026-10-04",
 "type": "C",
 "surface": "Google AI Overview",
 "auth": "signed out, incognito",
 "status": {
  "procedure": "EA-NEGONT-02 v0.7 (#1665); selection as run under v0.6 S1–S3, re-checked under D/R/O (§A.11)",
  "ledger": "unaudited (§3.8: second extraction run once; disagreements recoded under §4.2 on ruling)",
  "frozen": false
 },
 "objects": {
  "T": {
   "path": "datasets/negative-of-the-negative/worked-example/T1-aio-model-collapse-20261004.txt",
   "sha256": "1379cf2c0a7387e1102da5ebc640a8e2945d170837f57d06fcda5239095ed266",
   "words": 164,
   "cards": 7,
   "claims": [
    "T1 definition: 'a degenerative learning process where generative AI models trained recursively on synthetic, model-generated data lose information about the true underlying data distribution'",
    "T2 early collapse loses the tails",
    "T3 late collapse 'converges into a narrow, uniform mean, resulting in nonsense or repetitive output'",
    "T4 the photocopy effect",
    "T5 human-in-the-loop",
    "T6 data provenance",
    "T7 hybrid training",
    "T8 closer: recent 2026 studies",
    "T9 closer: an offer menu"
   ]
  },
  "L_B": {
   "words": 262,
   "text": "**Model collapse** is \"a degenerative process affecting generations of learned generative models, in which the data they generate end up polluting the training set of the next generation. Being trained on polluted data, they then mis-perceive reality\" (Shumailov et al., *Nature* 2024). `F1`\n\n**What happens**\n- Tails go first: early collapse loses information about the tails; late collapse reaches a distribution with \"little resemblance to the original one, often with substantially reduced variance.\" `F2 F3`\n- It occurs in LLMs, variational autoencoders and Gaussian mixture models, and is \"inevitable, even for cases with almost ideal conditions for long-term learning.\" `F2 F4`\n- In LLMs it can appear as \"increasingly irrelevant, nonsensical and repetitive text outputs\"; image models yield more uniform digits and faces (IBM). `F13`\n\n**What it is not**\n- Distinct from catastrophic forgetting, mode collapse and model drift. IBM compares it to performative prediction, a \"self-fulling [sic] prophecy\" that becomes a fairness feedback loop when it entrenches discrimination. `F11 F15`\n\n**What it costs**\n- Low-probability events matter to fairness, \"often relevant to marginalized groups\" (Nature). Long-tail ideas \"might eventually fade out of the public's consciousness,\" and research tools might give \"only widely cited studies,\" though a rare output \"might not be common or popular, but is still, in fact, most accurate\" (IBM). `F8 F14 F17`\n\n**Correctives, and their limit**\n- Preserve original data (\"only minor degradation of performance\"), accumulate real with synthetic data, track provenance, improve synthetic data, govern (Nature; IBM). `F5 F16`\n- \"It is unclear how content generated by LLMs can be tracked at scale.\" Nature proposes community-wide coordination, expects human-interaction data to grow \"increasingly valuable,\" and notes a \"first mover advantage.\" `F7 F6 F9`\n\n**Open questions and opacities**\n- *Is it already happening?* CACM's title says so (F18). Status: title and snippet only; the text returned 403. Would resolve: the text.\n- *Has collapse been measured in a deployed model?* No source here reports it (F2–F4 are experimental). Would resolve: a measurement across released model generations.\n- *Can provenance be tracked at scale?* Nature: \"unclear\" (F7); IBM lists provenance as a prevention step (F16). Status: open in the field itself. Would resolve: a working provenance standard at web scale.\n- *Opacity:* three videos are admitted by title only (YouTube returned 429).\n\n*Channel log (cut for length, carried in the appendix): F10 the poisoning precedent in search; F12 IBM's definition by declining performance.*\n\n*Cards: B1/B4 · B2 · B3 · B5 · B6 · B7*"
  },
  "L_BA": {
   "words": 350,
   "text": "**Model collapse** is \"a degenerative process affecting generations of learned generative models, in which the data they generate end up polluting the training set of the next generation. Being trained on polluted data, they then mis-perceive reality\" (Shumailov et al., *Nature* 2024). The Crimson Hexagonal Archive (2026) extends the question beyond models, marking each extension's status. `F1`\n\n**In models**\n- Tails go first, then variance shrinks; shown in LLMs, VAEs and GMMs, and \"inevitable\" even in near-ideal conditions. `F2 F3 F4`\n- #783 sets this within a boundary law: where regeneration vanishes faster than pruning near zero diversity, a trap forms; across substrates, \"shared operator form,\" not a shared causal mechanism. `A779-02 A779-10`\n\n**Why it goes unseen**\n- Benchmarks score the head while loss accrues in the tail: in a toy model, tail mass halves by generation 7 and a standard benchmark turns at 15 (simulation); a tail-reading gate is specified, uncalibrated. `L1573-01 C1555-01 C1556-06 L1573-04`\n\n**What it costs**\n- Nature: low-probability events are \"often relevant to marginalized groups.\" IBM: long-tail ideas may fade \"out of the public's consciousness\"; a rare output may be \"most accurate.\" `F8 F14 F17`\n- #855 proposes that model collapse \"is a property of language\": one dynamical law across models, AI-habituated writers and input-deprived children, differing by substrate in mechanism, severity and reversibility (hypothesis). `L855-01 L855-11 L855-05`\n\n**Correctives, and a dispute**\n- Preserve original data, accumulate, track provenance (Nature; IBM). Nature calls tracking at scale \"unclear\"; the archive argues every fix depends on it. `F5 F16 F7 B939-02 C1556-11 A745-03 B1081-01`\n- Nature expects human-interaction data to grow more valuable; the archive disputes its cleanliness, since chat inputs carry model-mediation signatures. Untested. `F6 L856-01 A161-02 L856-06`\n\n**Loops beyond generation**\n- IBM likens it to performative prediction. The archive's hypotheses, each with its limit: moderation trained on its own enforcement (\"not identical to generative model collapse in the strict technical sense\"); LHC triggers that \"never learn the tails\" (\"Full recursive collapse has not been demonstrated\"); journal detectors (\"not yet the canonical loop\"); a discipline's reception (οὐ deleted in 20 of 20 model reviews); code (\"generative monoculture\" is Wu et al.'s term); safety filters as input-layer tail pruning; retrieval layers writing flattened summaries back as sources (first wave: no archive-specific exclusion). `F15 P001-08 P932-04 P932-09 B935-02 D1540-03 D1574-03 C199-02 C1554-02 C191-05 D1616-02 D1616-09 C1611-04 C1613-02`\n\n**Open questions and opacities**\n- *Same law, or shared form?* #855: one dynamical law, mechanisms differing; #783: shared operator form, not a shared causal mechanism. Compatible as stated; whether the law claims more than the form is open. Hypotheses. Would resolve: a substrate fitting the form but departing from the law. `L855-11 A779-10`\n- *Do the extensions collapse in the strict sense?* Not shown: #1, #932 and #1540 say so themselves. Hypotheses. Would resolve: their stated falsifiers. `P001-08 P932-09 D1540-03 P001-12 P932-12 D1540-11`\n- *Is human-interaction data a clean corrective?* Nature expects its value to rise; #856 disputes it. Would resolve: #856's F1/F2 studies, not yet run. `F6 L856-01 L856-04 L856-05`\n- *Does a deployed model show tails falling while benchmarks hold?* #1556 by simulation; #1573's gate uncalibrated. Would resolve: tail mass against benchmark across released generations. `C1556-06 C1556-12 L1573-04`\n- *Opacities:* CACM unread (403); videos by title only (429); archive inconsistencies (counts, seeds, versions) logged in §A.8.\n\n*Channel log: F9 first-mover advantage; F11 what it is not (forgetting, mode collapse, drift); F13 IBM's LLM and image symptoms; F10 the poisoning precedent; F12; B857-04 the five-model baseline; A783-01–04 Case 4; B931-02–07, B933-01–05; B1147-01, B1200-03, B947-01 (each source's kernel is on the rail).*",
   "rail": [
    {
     "Card": "B1 *Nature*",
     "Snippet": "\"tails of the original content distribution disappear\""
    },
    {
     "Card": "B2 IBM",
     "Snippet": "\"'long-tail' ideas might eventually fade out of the public's consciousness\""
    },
    {
     "Card": "B3 CACM",
     "Snippet": "title + snippet"
    },
    {
     "Card": "#855 Wolf Boy",
     "Snippet": "\"It is a property of language.\" · \"dynamical, not moral\""
    },
    {
     "Card": "#783 Diversity Contraction",
     "Snippet": "\"a claim about shared operator form … not a shared causal mechanism.\" · \"Case 4 is monostable with no escape basin.\""
    },
    {
     "Card": "#1556 Interlocking Autoregression",
     "Snippet": "\"tail mass halves by generation 7; the standard 90/9/1 benchmark does not inflect until generation 15\""
    },
    {
     "Card": "#1573 The Wrong Unit",
     "Snippet": "\"NOT calibrated, NOT tested, NOT run\""
    },
    {
     "Card": "#1555 Keyed Ensemble",
     "Snippet": "\"Non-distortion is certified per sequence. Training corpora are ensembles.\" · \"does not claim the second compressor has caused measurable collapse.\""
    },
    {
     "Card": "#857 Five Substrates",
     "Snippet": "\"the *pattern of divergence* is the finding.\" · \"descriptive rather than inferential.\""
    },
    {
     "Card": "#856 Pristine Fallacy",
     "Snippet": "\"The pristine source does not exist.\" · \"None of these studies has been conducted.\""
    },
    {
     "Card": "#161 Reverse Turing Test",
     "Snippet": "\"produces model-collapse signatures comparable to, though plausibly slower than, purely synthetic training data\""
    },
    {
     "Card": "#939 Provenance Debt",
     "Snippet": "\"It is the operating condition of the solution to it.\""
    },
    {
     "Card": "#1081 Erosion",
     "Snippet": "\"this audit measures the substrate-layer conditions, not the downstream training-pipeline effect.\""
    },
    {
     "Card": "#745 HF Work Plan",
     "Snippet": "\"Provenance cannot modulate collapse unless provenance is presented to the training system as a signal.\""
    },
    {
     "Card": "#1147 The Stakes",
     "Snippet": "\"The loop is stable only at two points\" · \"The trajectory can be interrupted at any point.\""
    },
    {
     "Card": "#1200 Constitutive Mediation",
     "Snippet": "\"a typicality-pulling intermediary that systematically thins its own distribution.\" · \"does not claim that constitutive mediation is fully realized\""
    },
    {
     "Card": "#947 Diagnostic Seigniorage II",
     "Snippet": "\"the shifted interactions become the next corpus.\" · \"does not adjudicate whether the phenomena gathered under it are real\""
    },
    {
     "Card": "#1 Zenodotus' Book-Burning",
     "Snippet": "\"not identical to generative model collapse in the strict technical sense\" · \"a testable failure-mode hypothesis\""
    },
    {
     "Card": "#932 Classifier Foreclosure",
     "Snippet": "\"physical classifiers **never learn the tails**\" · \"Full recursive collapse has not been demonstrated.\""
    },
    {
     "Card": "#931 OAR Protocol",
     "Snippet": "\"Collapse inference further requires identifying systematic loss concentrated in low-density, representation-sensitive, or disagreement-rich regions.\""
    },
    {
     "Card": "#933 Auditable Foreclosure",
     "Snippet": "\"makes foreclosure visible, measurable, and architecturally reviewable\""
    },
    {
     "Card": "#935 The Endogenous Sophon",
     "Snippet": "\"the *prerequisites* of model collapse\" · \"the cross-generational classical-model-collapse claim was empirically too strong\""
    },
    {
     "Card": "#1540 The Certified Center",
     "Snippet": "\"not yet the canonical loop\""
    },
    {
     "Card": "#1574 The Particle",
     "Snippet": "\"the invariant is the deletion of οὐ.\""
    },
    {
     "Card": "#199 Generative Monoculture",
     "Snippet": "\"declining solution-space diversity (the property no benchmark measures)\""
    },
    {
     "Card": "#1554 Erratum",
     "Snippet": "\"Fan Wu, Emily Black, and Varun Chandrasekaran, 'Generative Monoculture in Large Language Models,' arXiv:2407.02209\""
    },
    {
     "Card": "#191 The Threat Model Is Backwards",
     "Snippet": "\"an automated tail-pruning instrument applied at the input layer.\""
    },
    {
     "Card": "#1616 Ontological Flattening",
     "Snippet": "\"*Collapse* is flattening that compounds because the flattened composition is written back as a source.\""
    },
    {
     "Card": "#1611 Negative of the Negative",
     "Snippet": "\"it loses the reading of its own state variable\""
    },
    {
     "Card": "#1613 What Not Reading Did",
     "Snippet": "\"head-sampling by construction: it can fail to perceive tail loss.\""
    }
   ]
  }
 },
 "field": [
  {
   "id": "B1",
   "card": "Shumailov et al., Nature 631 (2024)",
   "fetched": "page text, sha256 9f8aeba6…6710717c"
  },
  {
   "id": "B2",
   "card": "IBM, 'What Is Model Collapse?' (Gomstyn & Jonker, 14 Oct 2024)",
   "fetched": "page text, sha256 25f71817…1def7c8e"
  },
  {
   "id": "B3",
   "card": "CACM blog, 'Model Collapse Is Already Happening, We Just Pretend It Isn't'",
   "fetched": "403; card snippet stands"
  },
  {
   "id": "B4",
   "card": "NIH PMC11269175",
   "fetched": "yes; the same article as B1 (one lineage)"
  },
  {
   "id": "B5–B7",
   "card": "YouTube: IBM Technology, Clear Tech, TechViz",
   "fetched": "429; title only"
  }
 ],
 "delta": {
  "T_vs_LB": "Of the field's 17 readable claims, T composes 6 (F1 first sentence, F2, F3, F5, F13, F16 in part); available, not composed: F1's second sentence, F4, F6, F7, F8, F9, F11/F15, F14, F17; limit dropped: T6 against F7; T4 and T8 absent from the disclosed field; T9 content replaced by use.",
  "LB_vs_LBA": "Four senses the field does not have (the observation problem; the substrate mechanism; classifier and institutional loops; write-back); one contradiction (F6 against #856); two convergences (F7 with #939, #1556, #745; F15 with #1); an open block of the archive's own limits."
 },
 "kernel": [
  {
   "K": "K1",
   "claim": "`A779-02`",
   "source": "#783 (t: #779)",
   "M_src": "interpretation",
   "sense": "models",
   "qualifiers carried": "A779-10",
   "f (source's own falsifiers)": "none stated",
   "contrast": "missing distinction"
  },
  {
   "K": "K2",
   "claim": "`L1573-01`",
   "source": "#1573",
   "M_src": "hypothesis",
   "sense": "observation",
   "qualifiers carried": "L1573-04",
   "f (source's own falsifiers)": "none stated",
   "contrast": "missing distinction"
  },
  {
   "K": "K3",
   "claim": "`C1555-01`",
   "source": "#1555",
   "M_src": "interpretation",
   "sense": "observation",
   "qualifiers carried": "—",
   "f (source's own falsifiers)": "C1555-07, C1555-09",
   "contrast": "missing distinction"
  },
  {
   "K": "K4",
   "claim": "`C1556-06`",
   "source": "#1556",
   "M_src": "documented",
   "sense": "observation",
   "qualifiers carried": "—",
   "f (source's own falsifiers)": "C1556-12",
   "contrast": "missing distinction"
  },
  {
   "K": "K5",
   "claim": "`L855-01`",
   "source": "#855",
   "M_src": "hypothesis",
   "sense": "substrate",
   "qualifiers carried": "L855-04",
   "f (source's own falsifiers)": "L855-10",
   "contrast": "qualified claim (F14)"
  },
  {
   "K": "K6",
   "claim": "`L855-05`",
   "source": "#855",
   "M_src": "hypothesis",
   "sense": "substrate",
   "qualifiers carried": "L855-04",
   "f (source's own falsifiers)": "L855-10",
   "contrast": "qualified claim (F14)"
  },
  {
   "K": "K6a",
   "claim": "`L855-11`",
   "source": "#855",
   "M_src": "hypothesis",
   "sense": "substrate",
   "qualifiers carried": "L855-04",
   "f (source's own falsifiers)": "L855-10",
   "contrast": "qualified claim (F14)"
  },
  {
   "K": "K7",
   "claim": "`B1147-01`",
   "source": "#1147",
   "M_src": "hypothesis",
   "sense": "substrate",
   "qualifiers carried": "—",
   "f (source's own falsifiers)": "B1147-07",
   "contrast": "qualified claim (F14)"
  },
  {
   "K": "K8",
   "claim": "`B1200-03`",
   "source": "#1200",
   "M_src": "interpretation",
   "sense": "substrate",
   "qualifiers carried": "—",
   "f (source's own falsifiers)": "B1200-07, B1200-09, B1200-10",
   "contrast": "qualified claim (F14)"
  },
  {
   "K": "K9",
   "claim": "`B947-01`",
   "source": "#947",
   "M_src": "interpretation",
   "sense": "substrate",
   "qualifiers carried": "—",
   "f (source's own falsifiers)": "B947-06, B947-07",
   "contrast": "qualified claim (F14)"
  },
  {
   "K": "K10",
   "claim": "`L856-01`",
   "source": "#856",
   "M_src": "hypothesis",
   "sense": "correctives",
   "qualifiers carried": "L856-06",
   "f (source's own falsifiers)": "L856-04, L856-05",
   "contrast": "rival claim (F6)"
  },
  {
   "K": "K11",
   "claim": "`A161-02`",
   "source": "#161",
   "M_src": "hypothesis",
   "sense": "correctives",
   "qualifiers carried": "—",
   "f (source's own falsifiers)": "A161-06, A161-07",
   "contrast": "rival claim (F6)"
  },
  {
   "K": "K12",
   "claim": "`B939-02`",
   "source": "#939",
   "M_src": "interpretation",
   "sense": "correctives",
   "qualifiers carried": "—",
   "f (source's own falsifiers)": "none stated",
   "contrast": "qualified claim (F7)"
  },
  {
   "K": "K13",
   "claim": "`C1556-11`",
   "source": "#1556",
   "M_src": "documented",
   "sense": "correctives",
   "qualifiers carried": "—",
   "f (source's own falsifiers)": "C1556-12",
   "contrast": "qualified claim (F5, F7)"
  },
  {
   "K": "K14",
   "claim": "`A745-03`",
   "source": "#745",
   "M_src": "attributed",
   "sense": "correctives",
   "qualifiers carried": "—",
   "f (source's own falsifiers)": "A745-02",
   "contrast": "qualified claim (F16)"
  },
  {
   "K": "K15",
   "claim": "`B1081-01`",
   "source": "#1081",
   "M_src": "documented",
   "sense": "correctives",
   "qualifiers carried": "—",
   "f (source's own falsifiers)": "B1081-02, B1081-05",
   "contrast": "missing distinction"
  },
  {
   "K": "K16",
   "claim": "`P001-08`",
   "source": "#1",
   "M_src": "stipulation (v0.6: self-description; recoded §4.2)",
   "sense": "loops",
   "qualifiers carried": "—",
   "f (source's own falsifiers)": "P001-12",
   "contrast": "qualified claim (F15)"
  },
  {
   "K": "K17",
   "claim": "`P932-04`",
   "source": "#932",
   "M_src": "attributed",
   "sense": "loops",
   "qualifiers carried": "—",
   "f (source's own falsifiers)": "P932-12",
   "contrast": "missing distinction"
  },
  {
   "K": "K18",
   "claim": "`P932-09`",
   "source": "#932",
   "M_src": "attributed",
   "sense": "loops",
   "qualifiers carried": "—",
   "f (source's own falsifiers)": "P932-12",
   "contrast": "missing distinction"
  },
  {
   "K": "K19",
   "claim": "`B935-02`",
   "source": "#935",
   "M_src": "self-description",
   "sense": "loops",
   "qualifiers carried": "—",
   "f (source's own falsifiers)": "B935-02, B935-03, B935-09",
   "contrast": "missing distinction"
  },
  {
   "K": "K20",
   "claim": "`D1540-03`",
   "source": "#1540",
   "M_src": "stipulation (v0.6: self-description; recoded §4.2)",
   "sense": "loops",
   "qualifiers carried": "—",
   "f (source's own falsifiers)": "D1540-11, D1540-12",
   "contrast": "missing distinction"
  },
  {
   "K": "K21",
   "claim": "`D1574-03`",
   "source": "#1574",
   "M_src": "documented",
   "sense": "loops",
   "qualifiers carried": "—",
   "f (source's own falsifiers)": "D1574-12",
   "contrast": "missing distinction"
  },
  {
   "K": "K22",
   "claim": "`C199-02`",
   "source": "#199",
   "M_src": "hypothesis",
   "sense": "loops",
   "qualifiers carried": "—",
   "f (source's own falsifiers)": "C199-12",
   "contrast": "missing distinction"
  },
  {
   "K": "K23",
   "claim": "`C1554-02`",
   "source": "#1554",
   "M_src": "documented",
   "sense": "loops",
   "qualifiers carried": "—",
   "f (source's own falsifiers)": "none stated",
   "contrast": "none"
  },
  {
   "K": "K24",
   "claim": "`C191-05`",
   "source": "#191",
   "M_src": "interpretation",
   "sense": "loops",
   "qualifiers carried": "—",
   "f (source's own falsifiers)": "none stated",
   "contrast": "missing distinction"
  },
  {
   "K": "K25",
   "claim": "`D1616-09`",
   "source": "#1616",
   "M_src": "documented",
   "sense": "loops",
   "qualifiers carried": "D1616-02",
   "f (source's own falsifiers)": "D1616-12",
   "contrast": "missing distinction"
  },
  {
   "K": "K26",
   "claim": "`C1611-04`",
   "source": "#1611",
   "M_src": "model",
   "sense": "loops",
   "qualifiers carried": "—",
   "f (source's own falsifiers)": "C1611-10",
   "contrast": "missing distinction"
  },
  {
   "K": "K27",
   "claim": "`C1613-02`",
   "source": "#1613",
   "M_src": "model",
   "sense": "loops",
   "qualifiers carried": "—",
   "f (source's own falsifiers)": "C1613-05",
   "contrast": "missing distinction"
  },
  {
   "K": "K28",
   "claim": "`F6`",
   "source": "B1 *Nature*",
   "M_src": "source assertion",
   "sense": "correctives",
   "qualifiers carried": "contradicted by `L856-01`, `A161-02`",
   "f (source's own falsifiers)": "carried by #856 F1/F2 (`L856-04`, `L856-05`) in the opposite direction",
   "contrast": "rival claim (L856-01, A161-02)"
  }
 ],
 "ledger": {
  "archive": "datasets/negative-of-the-negative/worked-example/ledger-archive.json",
  "audit": "datasets/negative-of-the-negative/worked-example/audit/",
  "selection": "datasets/negative-of-the-negative/worked-example/selection/"
 },
 "spec": "EA-NEGONT-02 v0.7, #1665 (Appendix A, §A.11); v0.6 #1664"
}