The Negative of the Negative › entity model-collapse

One entity of the Negative of the Negative (EA-NEGONT-02, #1665), cited at https://www.alexanarch.org/non/model-collapse/. the table of contents · this entity as data · contents as data · row json · archive ledger · D/R/O traversal · the address page.

model collapse

type C — a conventional reading holds it · sought at model collapse

worked example: T, P, C, D, K drafted (Appendix A)

Google AI Overviewepoch 2026-10-04
not frozentype C — a conventional reading holds itsigned out, incognitoknowledge object · 15 sentences

EA-NEGONT-02 v0.7 (#1665); selection as run under v0.6 S1–S3, re-checked under D/R/O (§A.11). Ledger: unaudited (§3.8: second extraction run once; disagreements recoded under §4.2 on ruling).

WORKING KNOWLEDGE OBJECT · NOT FROZEN · LEDGER UNAUDITED

Compression · Field and archive (B ∪ A)model collapse

Model collapse is the progressive loss of rare distinctions when a system learns from outputs it helped produce. It is demonstrated in generative AI; analogous collapse in other selection systems is described, and one general law is proposed, though not established.

🔁 How it unfolds

  • Tails first. Rare cases go; the output narrows toward a low-variance copy of the original.2↘
    Information about the tails of the original distribution is lost first; later, the learned distribution converges toward one with little resemblance to the original and much reduced variance.§2 The effect has been shown in large language models, variational autoencoders and Gaussian mixture models, and under indiscriminate recursive training it is inevitable even in conditions close to ideal.§3
  • Silent. The common survives longest, so benchmarks can hold while the tail disappears.2↘
    Because the head of a distribution survives longest, collapse can proceed while standard evaluations hold steady: in a toy model, tail mass halves by the seventh generation while a standard benchmark does not turn until the fifteenth.§6 A diagnostic that samples only what a system already admits is head-sampling by construction and can fail to perceive tail loss; an instrument that reads the tail directly has been specified but not calibrated, tested or run.§7
  • What goes first: minority cases, long-tail ideas, the accurate answer that is not the popular one.1↘
    What is lost is the rare: low-probability events are often those relevant to marginalized groups, long-tail ideas may fade from public consciousness, and a rare output, though neither common nor popular, may be the accurate one.§8

📍 Where it appears

  • Demonstrated: language models, autoencoders, mixture models.3↘
    Information about the tails of the original distribution is lost first; later, the learned distribution converges toward one with little resemblance to the original and much reduced variance.§2 The effect has been shown in large language models, variational autoencoders and Gaussian mixture models, and under indiscriminate recursive training it is inevitable even in conditions close to ideal.§3 In language models it can appear as increasingly irrelevant, nonsensical or repetitive text; in image models, as digits and faces that grow more alike.§4
  • Analogous mechanisms described: moderation classifiers trained on their own enforcement · physics triggers · journal screening · code (solution diversity falls while correctness holds) · AI-text detectors · search summaries written back as sources.2↘
    Recursive narrowing of the same shape has been described outside generative training: in moderation classifiers trained on their own enforcement, in particle-physics triggers that never learn the tails of the physical distribution, in journal screening, in the reception of a discipline, in code, where it shows as declining solution-space diversity rather than declining correctness, in detectors that prune high-perplexity input, and in retrieval layers whose flattened summaries are written back as sources.§14 In none of these has full recursive collapse in the strict technical sense been demonstrated.§15
  • General law proposed: one dynamic across models, AI-habituated writers and language-deprived children, differing in whether the loss reverses. A cautious reading claims only a shared form.2↘
    On one proposal, model collapse is a property of language rather than of language models: a single dynamical law would govern recursively trained models, writers habituated to AI-generated text and children deprived of linguistic input, the three differing in severity, mechanism and timescale, and in whether the loss can be reversed.§9 Described more cautiously, what such cases share is an operator form, transmission composed with selection, and not a common causal mechanism.§10

🛡️ Prevention

  • Keep the original data; accumulate real data alongside synthetic.1↘
    Preserving original data keeps degradation minor, and accumulating real data alongside synthetic, determining provenance, improving synthetic data and governance tools are proposed as preventives.§11
  • Provenance works only if the training system receives it as a signal.1↘
    How model-generated content can be tracked at scale is unclear, and the preventives depend on it: on one analysis provenance is the operating condition of any solution, capping a corpus's synthetic share presupposes telling synthetic from human text, and provenance cannot modulate collapse unless it reaches the training system as a signal.§12
  • Human data grows more valuable, though chat-era data may already carry model signatures (untested).1↘
    Data from genuine human interaction is expected to grow more valuable; whether it is a clean corrective is disputed, since human inputs to chat systems may already carry the signatures of model mediation and training on them may produce collapse signatures comparable to those of synthetic data, if more slowly, and the studies that would decide it have not been conducted.§13
1Recursive synthetic-data collapseShumailov et al., "AI models collapse when trained on recursively generated data", Nature 631 (2024)5 claims · B1, B2
  • F1 Shumailov et al., "AI models collapse when trained on recursively generated data", Nature 631 (2024), Def. 2.1 documented · field
    Model collapse is a degenerative process affecting generations of learned generative models, in which the data they generate end up polluting the training set of the next generation. Being trained on polluted data, they then mis-perceive reality.
  • F2 Shumailov et al., "AI models collapse when trained on recursively generated data", Nature 631 (2024), Abstract documented · field
    indiscriminate use of model-generated content in training causes irreversible defects in the resulting models, in which tails of the original content distribution disappear … it can occur in LLMs as well as in variational autoencoders (VAEs) and Gaussian mixture models (GMMs)
  • F3 Shumailov et al., "AI models collapse when trained on recursively generated data", Nature 631 (2024), text documented · field
    early collapse loses "information about the tails"; late collapse converges on "a distribution that carries little resemblance to the original one, often with substantially reduced variance"
  • F4 Shumailov et al., "AI models collapse when trained on recursively generated data", Nature 631 (2024), Main documented · field
    this process is inevitable, even for cases with almost ideal conditions for long-term learning
  • F13 IBM, "What Is Model Collapse?" (Gomstyn & Jonker, 14 Oct 2024), text documented · field
    In LLMs, model collapse can manifest in increasingly irrelevant, nonsensical and repetitive text outputs"; image models give digits that resemble each other and "more homogeneous faces
2What collapse forgets is the consequential rareShumailov et al., "AI models collapse when trained on recursively generated data", Nature 631 (2024)3 claims · B1, B2
  • F8 Shumailov et al., "AI models collapse when trained on recursively generated data", Nature 631 (2024), Discussion documented · field
    Preserving the ability of LLMs to model low-probability events is essential to the fairness of their predictions: such events are often relevant to marginalized groups.
  • F14 IBM, "What Is Model Collapse?" (Gomstyn & Jonker, 14 Oct 2024), text documented · field
    consequences: poor decision-making (a rare disease "forgotten"); user disengagement; knowledge decline, "'long-tail' ideas might eventually fade out of the public's consciousness"; research tools "might provide only widely cited studies"
  • F17 IBM, "What Is Model Collapse?" (Gomstyn & Jonker, 14 Oct 2024), text documented · field
    a rare output "might not be common or popular, but is still, in fact, most accurate" (the "rarely cited study")
3Preserving and accumulating real dataShumailov et al., "AI models collapse when trained on recursively generated data", Nature 631 (2024)2 claims · B1, B2
  • F5 Shumailov et al., "AI models collapse when trained on recursively generated data", Nature 631 (2024), text documented · field
    preservation of the original data allows for better model fine-tuning and leads to only minor degradation of performance.
  • F16 IBM, "What Is Model Collapse?" (Gomstyn & Jonker, 14 Oct 2024), text documented · field
    prevention: retaining non-AI data sources; determining data provenance (the Data Provenance Initiative, "more than 4,000 datasets"); data accumulation; better synthetic data; governance tools
4Silent tail loss: the head survives, benchmarks hold#1556 The Interlocking Autoregression: Three Coupled Recursions Under a Mism · 2026-08-273 claims · #1556, #1573, #1613
  • C1556-06 #1556 The Interlocking Autoregression: Three Coupled Recursions Under a Mismatched Observation Regime, with Toy Dynamics (EA-L Nobel Glas, Director, Lagrange Observatory! (LO!) · 2026-08-27 · §6.2 F1 documented (toy simulation, seed 20260827)
    Full loop: tail mass halves by generation 7; the standard 90/9/1 benchmark does not inflect until generation 15; head-only probes, generation 17.
  • L1573-01 #1573 The Wrong Unit: A Model-Collapse Self-Diagnostic in Three Grades — for Benchmarking, for Frontier Models, and for the Re Nobel Glas, Director, Lagrange Observatory! (LO!) · 2026-09-01 · §I.2 hypothesis
    each item is answered from the *head* of the model's distribution
  • C1613-02 #1613 What Not Reading Did to Its Own Ontology: The Mirror, Held Up — the damage a source-admission ontology does to itself wh Sharks, Lee · 2026-09-15 · §6 model consequence (hypothesis, "in the modelled regime")
    a diagnostic that samples the admitted component is head-sampling by construction: it can fail to perceive tail loss.
5Tail-pruning selection beyond generation#191 The Threat Model Is Backwards: On Classifying High-Perplexity Text as · 2026-06-116 claims · #1, #1540, #1574, #191, #932
  • C191-05 #191 The Threat Model Is Backwards: On Classifying High-Perplexity Text as a Security Threat in an Era of Model Collapse Lee Sharks (primary), with Nobel Glas and Talos Morrow · 2026-06-11 · §3 interpretation (self-typed *Structural*)
    A high-perplexity-content detector that rejects on positive detection is, viewed from the model-collapse literature, an automated tail-pruning instrument applied at the input layer.
  • P001-08 #1 Zenodotus' Book-Burning: Loud Exclusion at Repository Scale Lee Sharks · 2026-06-19 · §6 stipulation
    Classifier model collapse, as defined here, is not identical to generative model collapse in the strict technical sense (Shumailov et al., 2024). It names a moderation-specific feedback contraction:
  • P932-04 #932 EA-SEI-COLLAPSE-SYNTHESIS-01 v0.3: Classifier Foreclosure in Physical Measurement — Substrate Witnesses, Integrative Syn Lee Sharks · 2026-06-29 · Appendix W1, §10. Distinction from Generative Model Collapse attributed (Witness 1)
    Classifier collapse is the **discriminative analogue** of generative model collapse. Where Shumailov's models forget the tails of their own distribution, physical classifiers **never learn the tails** of the true physical distribution. The collapse is present from the first forward pass.
  • P932-09 #932 EA-SEI-COLLAPSE-SYNTHESIS-01 v0.3: Classifier Foreclosure in Physical Measurement — Substrate Witnesses, Integrative Syn Lee Sharks · 2026-06-29 · Appendix W3, §4. Are the operations of full collapse already visible?; §1.3 attributed (Witness 3), adopted by the synthesis
    **The ingredients are visible. Full recursive collapse has not been demonstrated.**" / "*the LHC community has built an architecture in which phenomenal model collapse is possible, and the current validation literature does not yet demonstrate that it has been ruled out.*
  • D1540-03 #1540 The Certified Center: Retroactive Classifier Standing and the Institutional Path to Model Collapse in Philosophy Johannes Sigil; Nobel Glas · 2026-08-24 · §IV stipulation
    This stage alone establishes recursive distributional contraction under a statistical boundary — classifier-mediated recursive selection over a discourse — but not yet the canonical loop, which requires filtered outputs to recur into later model training.
  • D1574-03 #1574 The Particle: Provenance Erasure at Sophistical Refutations 183b34, Measured — Machine-Mediated Reception, Disciplinary Lee Sharks · 2026-09-02 · §3.3 The terminal form documented
    The operation is therefore measured to one particle. Twenty rounds, two models, eleven configurations; the invariant is the deletion of οὐ.
6Code: diversity before correctness#199 Generative Monoculture Model Collapse in Code as Systemic Vulnerabilit · 2026-06-131 claim · #199
  • C199-02 #199 Generative Monoculture Model Collapse in Code as Systemic Vulnerability Talos Morrow · Nobel Glas; contributing editor Lee Sharks · 2026-06-13 · Abstract ¶2, first claim hypothesis
    model collapse in code does not manifest primarily as declining functional correctness (the property benchmarks measure) but as declining solution-space diversity (the property no benchmark measures)
7One law across substrates (proposed)#855 The Wolf Boy and the Language Model: Model Collapse as Substrate-Agnos · 2026-06-183 claims · #855
  • L855-01 #855 The Wolf Boy and the Language Model: Model Collapse as Substrate-Agnostic Capacity Loss Nobel Glas · 2026-06-18 · Abstract hypothesis
    This paper argues it is not a property of language models. It is a property of language.
  • L855-11 #855 The Wolf Boy and the Language Model: Model Collapse as Substrate-Agnostic Capacity Loss Nobel Glas · 2026-06-18 · §II The Structural Identity hypothesis
    The three cases differ in severity, in mechanism, and in timescale. But they are governed by the same dynamical law
  • L855-05 #855 The Wolf Boy and the Language Model: Model Collapse as Substrate-Agnostic Capacity Loss Nobel Glas · 2026-06-18 · §II The Structural Identity hypothesis
    The dynamical law is the same; the reversibility parameter is substrate-specific.
8Shared form, cause unshown#779 Diversity Contraction Across Substrates: A Boundary Law for Semantic E · 2026-06-022 claims · #779
  • A779-10 #779 Diversity Contraction Across Substrates: A Boundary Law for Semantic Exhaustion Nobel Glas · 2026-06-02 · §4 The operator family self-description
    This is a claim about shared operator form — that each domain's dynamics can be written as transmission composed with selection — not a shared causal mechanism.
  • A779-02 #779 Diversity Contraction Across Substrates: A Boundary Law for Semantic Exhaustion Nobel Glas · 2026-06-02 · §4 The operator family (table) interpretation
    | Model collapse | Resampling from the model's own output distribution | Loss-minimization on a fixed objective | Endogenous (case 3) |
9Provenance as the operating conditionShumailov et al., "AI models collapse when trained on recursively generated data", Nature 631 (2024)4 claims · #1556, #745, #939, B1
  • F7 Shumailov et al., "AI models collapse when trained on recursively generated data", Nature 631 (2024), Discussion documented · field
    it is unclear how content generated by LLMs can be tracked at scale. One option is community-wide coordination
  • B939-02 #939 EA-PROVENANCE-DEBT-01 v0.2: Provenance Debt and the Extraction Economy of Unmarked Augmentation Lee Sharks · 2026-07-01 · §2 l.91 interpretation
    This means provenance is not adjacent to the model collapse question. It is the operating condition of the solution to it.
  • A745-03 #745 Crimson Hexagonal Archive — Hugging Face Dataset Work Plan v3 Lee Sharks · 2026-05-20 · Research Question, Operationalized attributed (to "Assembly review")
    Provenance cannot modulate collapse unless provenance is presented to the training system as a signal.
  • C1556-11 #1556 The Interlocking Autoregression: Three Coupled Recursions Under a Mismatched Observation Regime, with Toy Dynamics (EA-L Nobel Glas, Director, Lagrange Observatory! (LO!) · 2026-08-27 · §6.2 F5 documented (toy) + interpretation
    The Shumailov mitigation reproduces in the interlocked setting — and the lever it requires is the distinction Component I withholds *from downstream builders*: capping synthetic share presupposes the ability to distinguish it.
10Human data grows in valueShumailov et al., "AI models collapse when trained on recursively generated data", Nature 631 (2024)1 claim · B1
  • F6 Shumailov et al., "AI models collapse when trained on recursively generated data", Nature 631 (2024), Abstract documented · field
    the value of data collected about genuine human interactions with systems will be increasingly valuable
11Chat data is not pristine#161 The Reverse Turing Test: A Three-Stage Protocol for Detecting AI-Media · 2026-06-073 claims · #161, #856
  • A161-02 #161 The Reverse Turing Test: A Three-Stage Protocol for Detecting AI-Mediation Signatures in Human Text and Their Propagatio Lee Sharks · 2026-06-07 · Abstract hypothesis
    that training on AI-mediated human text — including unaided text from cognitively-habituated writers — produces model-collapse signatures comparable to, though plausibly slower than, purely synthetic training data
  • L856-01 #856 The Pristine Fallacy: Why Chat Data Is Not a Clean Training Source Lee Sharks · 2026-06-18 · Abstract hypothesis
    Chat data fails the pristine test on three independent grounds.
  • L856-06 #856 The Pristine Fallacy: Why Chat Data Is Not a Clean Training Source Lee Sharks · 2026-06-18 · §VI self-description
    None of these studies has been conducted.
12Flattening written back as source#1616 Ontological Flattening: Toy Models of the Collapse of Distinctions in · 2026-09-151 claim · #1616
  • D1616-02 #1616 Ontological Flattening: Toy Models of the Collapse of Distinctions in a Represented World, the Instrument They Specified Sharks, Lee · 2026-09-15 · §0 stipulation
    *Flattening* is the loss of a distinction from the reachable set while the things distinguished still exist. *Collapse* is flattening that compounds because the flattened composition is written back as a source.
Expansion · Field and archive (B ∪ A)15 sentences · every one sourced below

Model collapse

Model collapse is a degenerative process in generative models trained, generation after generation, on data that earlier models produced: the generated data pollute the next training set, and models trained on it come to misperceive the reality they were built to model. Information about the tails of the original distribution is lost first; later, the learned distribution converges toward one with little resemblance to the original and much reduced variance. The effect has been shown in large language models, variational autoencoders and Gaussian mixture models, and under indiscriminate recursive training it is inevitable even in conditions close to ideal. In language models it can appear as increasingly irrelevant, nonsensical or repetitive text; in image models, as digits and faces that grow more alike. It is distinct from catastrophic forgetting, mode collapse and model drift, and close to performative prediction, a self-fulfilling loop that becomes a fairness feedback loop when it entrenches discrimination.

Because the head of a distribution survives longest, collapse can proceed while standard evaluations hold steady: in a toy model, tail mass halves by the seventh generation while a standard benchmark does not turn until the fifteenth. A diagnostic that samples only what a system already admits is head-sampling by construction and can fail to perceive tail loss; an instrument that reads the tail directly has been specified but not calibrated, tested or run. What is lost is the rare: low-probability events are often those relevant to marginalized groups, long-tail ideas may fade from public consciousness, and a rare output, though neither common nor popular, may be the accurate one.

On one proposal, model collapse is a property of language rather than of language models: a single dynamical law would govern recursively trained models, writers habituated to AI-generated text and children deprived of linguistic input, the three differing in severity, mechanism and timescale, and in whether the loss can be reversed. Described more cautiously, what such cases share is an operator form, transmission composed with selection, and not a common causal mechanism.

Preserving original data keeps degradation minor, and accumulating real data alongside synthetic, determining provenance, improving synthetic data and governance tools are proposed as preventives. How model-generated content can be tracked at scale is unclear, and the preventives depend on it: on one analysis provenance is the operating condition of any solution, capping a corpus's synthetic share presupposes telling synthetic from human text, and provenance cannot modulate collapse unless it reaches the training system as a signal. Data from genuine human interaction is expected to grow more valuable; whether it is a clean corrective is disputed, since human inputs to chat systems may already carry the signatures of model mediation and training on them may produce collapse signatures comparable to those of synthetic data, if more slowly, and the studies that would decide it have not been conducted.

Recursive narrowing of the same shape has been described outside generative training: in moderation classifiers trained on their own enforcement, in particle-physics triggers that never learn the tails of the physical distribution, in journal screening, in the reception of a discipline, in code, where it shows as declining solution-space diversity rather than declining correctness, in detectors that prune high-perplexity input, and in retrieval layers whose flattened summaries are written back as sources. In none of these has full recursive collapse in the strict technical sense been demonstrated.

Provenance — every sentence sourced15 sentences

15 sentences; 14 field claims, 23 archive claims. Modality is the source's own; the prose carries it in its grammar.

1.Model collapse is a degenerative process in generative models trained, generation after generation, on data that earlier models produced: the generated data pollute the next training set, and models trained on it come to misperceive the reality they were built to model.1
  • F1 Shumailov et al., "AI models collapse when trained on recursively generated data", Nature 631 (2024), Def. 2.1 documented · field
    Model collapse is a degenerative process affecting generations of learned generative models, in which the data they generate end up polluting the training set of the next generation. Being trained on polluted data, they then mis-perceive reality.
2.Information about the tails of the original distribution is lost first; later, the learned distribution converges toward one with little resemblance to the original and much reduced variance.2
  • F2 Shumailov et al., "AI models collapse when trained on recursively generated data", Nature 631 (2024), Abstract documented · field
    indiscriminate use of model-generated content in training causes irreversible defects in the resulting models, in which tails of the original content distribution disappear … it can occur in LLMs as well as in variational autoencoders (VAEs) and Gaussian mixture models (GMMs)
  • F3 Shumailov et al., "AI models collapse when trained on recursively generated data", Nature 631 (2024), text documented · field
    early collapse loses "information about the tails"; late collapse converges on "a distribution that carries little resemblance to the original one, often with substantially reduced variance"
3.The effect has been shown in large language models, variational autoencoders and Gaussian mixture models, and under indiscriminate recursive training it is inevitable even in conditions close to ideal.2
  • F2 Shumailov et al., "AI models collapse when trained on recursively generated data", Nature 631 (2024), Abstract documented · field
    indiscriminate use of model-generated content in training causes irreversible defects in the resulting models, in which tails of the original content distribution disappear … it can occur in LLMs as well as in variational autoencoders (VAEs) and Gaussian mixture models (GMMs)
  • F4 Shumailov et al., "AI models collapse when trained on recursively generated data", Nature 631 (2024), Main documented · field
    this process is inevitable, even for cases with almost ideal conditions for long-term learning
4.In language models it can appear as increasingly irrelevant, nonsensical or repetitive text; in image models, as digits and faces that grow more alike.1
  • F13 IBM, "What Is Model Collapse?" (Gomstyn & Jonker, 14 Oct 2024), text documented · field
    In LLMs, model collapse can manifest in increasingly irrelevant, nonsensical and repetitive text outputs"; image models give digits that resemble each other and "more homogeneous faces
5.It is distinct from catastrophic forgetting, mode collapse and model drift, and close to performative prediction, a self-fulfilling loop that becomes a fairness feedback loop when it entrenches discrimination.2
  • F11 Shumailov et al., "AI models collapse when trained on recursively generated data", Nature 631 (2024), Main documented · field
    catastrophic forgetting and data poisoning are close concepts; "Neither is able to explain the phenomenon of model collapse fully"
  • F15 IBM, "What Is Model Collapse?" (Gomstyn & Jonker, 14 Oct 2024), text documented · field
    distinct from catastrophic forgetting, mode collapse and model drift; compared to performative prediction, a "self-fulling [sic] prophecy", "also known as a fairness feedback loop when this process entrenches discrimination"
6.Because the head of a distribution survives longest, collapse can proceed while standard evaluations hold steady: in a toy model, tail mass halves by the seventh generation while a standard benchmark does not turn until the fifteenth.1
  • C1556-06 #1556 The Interlocking Autoregression: Three Coupled Recursions Under a Mismatched Observation Regime, with Toy Dynamics (EA-L Nobel Glas, Director, Lagrange Observatory! (LO!) · 2026-08-27 · §6.2 F1 documented (toy simulation, seed 20260827)
    Full loop: tail mass halves by generation 7; the standard 90/9/1 benchmark does not inflect until generation 15; head-only probes, generation 17.
7.A diagnostic that samples only what a system already admits is head-sampling by construction and can fail to perceive tail loss; an instrument that reads the tail directly has been specified but not calibrated, tested or run.3
  • C1613-02 #1613 What Not Reading Did to Its Own Ontology: The Mirror, Held Up — the damage a source-admission ontology does to itself wh Sharks, Lee · 2026-09-15 · §6 model consequence (hypothesis, "in the modelled regime")
    a diagnostic that samples the admitted component is head-sampling by construction: it can fail to perceive tail loss.
  • L1573-01 #1573 The Wrong Unit: A Model-Collapse Self-Diagnostic in Three Grades — for Benchmarking, for Frontier Models, and for the Re Nobel Glas, Director, Lagrange Observatory! (LO!) · 2026-09-01 · §I.2 hypothesis
    each item is answered from the *head* of the model's distribution
  • L1573-04 #1573 The Wrong Unit: A Model-Collapse Self-Diagnostic in Three Grades — for Benchmarking, for Frontier Models, and for the Re Nobel Glas, Director, Lagrange Observatory! (LO!) · 2026-09-01 · front matter self-description
    The instrument is specified and NOT calibrated, NOT tested, NOT run
8.What is lost is the rare: low-probability events are often those relevant to marginalized groups, long-tail ideas may fade from public consciousness, and a rare output, though neither common nor popular, may be the accurate one.3
  • F8 Shumailov et al., "AI models collapse when trained on recursively generated data", Nature 631 (2024), Discussion documented · field
    Preserving the ability of LLMs to model low-probability events is essential to the fairness of their predictions: such events are often relevant to marginalized groups.
  • F14 IBM, "What Is Model Collapse?" (Gomstyn & Jonker, 14 Oct 2024), text documented · field
    consequences: poor decision-making (a rare disease "forgotten"); user disengagement; knowledge decline, "'long-tail' ideas might eventually fade out of the public's consciousness"; research tools "might provide only widely cited studies"
  • F17 IBM, "What Is Model Collapse?" (Gomstyn & Jonker, 14 Oct 2024), text documented · field
    a rare output "might not be common or popular, but is still, in fact, most accurate" (the "rarely cited study")
9.On one proposal, model collapse is a property of language rather than of language models: a single dynamical law would govern recursively trained models, writers habituated to AI-generated text and children deprived of linguistic input, the three differing in severity, mechanism and timescale, and in whether the loss can be reversed.3
  • L855-01 #855 The Wolf Boy and the Language Model: Model Collapse as Substrate-Agnostic Capacity Loss Nobel Glas · 2026-06-18 · Abstract hypothesis
    This paper argues it is not a property of language models. It is a property of language.
  • L855-11 #855 The Wolf Boy and the Language Model: Model Collapse as Substrate-Agnostic Capacity Loss Nobel Glas · 2026-06-18 · §II The Structural Identity hypothesis
    The three cases differ in severity, in mechanism, and in timescale. But they are governed by the same dynamical law
  • L855-05 #855 The Wolf Boy and the Language Model: Model Collapse as Substrate-Agnostic Capacity Loss Nobel Glas · 2026-06-18 · §II The Structural Identity hypothesis
    The dynamical law is the same; the reversibility parameter is substrate-specific.
10.Described more cautiously, what such cases share is an operator form, transmission composed with selection, and not a common causal mechanism.2
  • A779-10 #779 Diversity Contraction Across Substrates: A Boundary Law for Semantic Exhaustion Nobel Glas · 2026-06-02 · §4 The operator family self-description
    This is a claim about shared operator form — that each domain's dynamics can be written as transmission composed with selection — not a shared causal mechanism.
  • A779-02 #779 Diversity Contraction Across Substrates: A Boundary Law for Semantic Exhaustion Nobel Glas · 2026-06-02 · §4 The operator family (table) interpretation
    | Model collapse | Resampling from the model's own output distribution | Loss-minimization on a fixed objective | Endogenous (case 3) |
11.Preserving original data keeps degradation minor, and accumulating real data alongside synthetic, determining provenance, improving synthetic data and governance tools are proposed as preventives.2
  • F5 Shumailov et al., "AI models collapse when trained on recursively generated data", Nature 631 (2024), text documented · field
    preservation of the original data allows for better model fine-tuning and leads to only minor degradation of performance.
  • F16 IBM, "What Is Model Collapse?" (Gomstyn & Jonker, 14 Oct 2024), text documented · field
    prevention: retaining non-AI data sources; determining data provenance (the Data Provenance Initiative, "more than 4,000 datasets"); data accumulation; better synthetic data; governance tools
12.How model-generated content can be tracked at scale is unclear, and the preventives depend on it: on one analysis provenance is the operating condition of any solution, capping a corpus's synthetic share presupposes telling synthetic from human text, and provenance cannot modulate collapse unless it reaches the training system as a signal.4
  • F7 Shumailov et al., "AI models collapse when trained on recursively generated data", Nature 631 (2024), Discussion documented · field
    it is unclear how content generated by LLMs can be tracked at scale. One option is community-wide coordination
  • B939-02 #939 EA-PROVENANCE-DEBT-01 v0.2: Provenance Debt and the Extraction Economy of Unmarked Augmentation Lee Sharks · 2026-07-01 · §2 l.91 interpretation
    This means provenance is not adjacent to the model collapse question. It is the operating condition of the solution to it.
  • C1556-11 #1556 The Interlocking Autoregression: Three Coupled Recursions Under a Mismatched Observation Regime, with Toy Dynamics (EA-L Nobel Glas, Director, Lagrange Observatory! (LO!) · 2026-08-27 · §6.2 F5 documented (toy) + interpretation
    The Shumailov mitigation reproduces in the interlocked setting — and the lever it requires is the distinction Component I withholds *from downstream builders*: capping synthetic share presupposes the ability to distinguish it.
  • A745-03 #745 Crimson Hexagonal Archive — Hugging Face Dataset Work Plan v3 Lee Sharks · 2026-05-20 · Research Question, Operationalized attributed (to "Assembly review")
    Provenance cannot modulate collapse unless provenance is presented to the training system as a signal.
13.Data from genuine human interaction is expected to grow more valuable; whether it is a clean corrective is disputed, since human inputs to chat systems may already carry the signatures of model mediation and training on them may produce collapse signatures comparable to those of synthetic data, if more slowly, and the studies that would decide it have not been conducted.4
  • F6 Shumailov et al., "AI models collapse when trained on recursively generated data", Nature 631 (2024), Abstract documented · field
    the value of data collected about genuine human interactions with systems will be increasingly valuable
  • L856-01 #856 The Pristine Fallacy: Why Chat Data Is Not a Clean Training Source Lee Sharks · 2026-06-18 · Abstract hypothesis
    Chat data fails the pristine test on three independent grounds.
  • A161-02 #161 The Reverse Turing Test: A Three-Stage Protocol for Detecting AI-Mediation Signatures in Human Text and Their Propagatio Lee Sharks · 2026-06-07 · Abstract hypothesis
    that training on AI-mediated human text — including unaided text from cognitively-habituated writers — produces model-collapse signatures comparable to, though plausibly slower than, purely synthetic training data
  • L856-06 #856 The Pristine Fallacy: Why Chat Data Is Not a Clean Training Source Lee Sharks · 2026-06-18 · §VI self-description
    None of these studies has been conducted.
14.Recursive narrowing of the same shape has been described outside generative training: in moderation classifiers trained on their own enforcement, in particle-physics triggers that never learn the tails of the physical distribution, in journal screening, in the reception of a discipline, in code, where it shows as declining solution-space diversity rather than declining correctness, in detectors that prune high-perplexity input, and in retrieval layers whose flattened summaries are written back as sources.7
  • P001-08 #1 Zenodotus' Book-Burning: Loud Exclusion at Repository Scale Lee Sharks · 2026-06-19 · §6 stipulation
    Classifier model collapse, as defined here, is not identical to generative model collapse in the strict technical sense (Shumailov et al., 2024). It names a moderation-specific feedback contraction:
  • P932-04 #932 EA-SEI-COLLAPSE-SYNTHESIS-01 v0.3: Classifier Foreclosure in Physical Measurement — Substrate Witnesses, Integrative Syn Lee Sharks · 2026-06-29 · Appendix W1, §10. Distinction from Generative Model Collapse attributed (Witness 1)
    Classifier collapse is the **discriminative analogue** of generative model collapse. Where Shumailov's models forget the tails of their own distribution, physical classifiers **never learn the tails** of the true physical distribution. The collapse is present from the first forward pass.
  • D1540-03 #1540 The Certified Center: Retroactive Classifier Standing and the Institutional Path to Model Collapse in Philosophy Johannes Sigil; Nobel Glas · 2026-08-24 · §IV stipulation
    This stage alone establishes recursive distributional contraction under a statistical boundary — classifier-mediated recursive selection over a discourse — but not yet the canonical loop, which requires filtered outputs to recur into later model training.
  • D1574-03 #1574 The Particle: Provenance Erasure at Sophistical Refutations 183b34, Measured — Machine-Mediated Reception, Disciplinary Lee Sharks · 2026-09-02 · §3.3 The terminal form documented
    The operation is therefore measured to one particle. Twenty rounds, two models, eleven configurations; the invariant is the deletion of οὐ.
  • C199-02 #199 Generative Monoculture Model Collapse in Code as Systemic Vulnerability Talos Morrow · Nobel Glas; contributing editor Lee Sharks · 2026-06-13 · Abstract ¶2, first claim hypothesis
    model collapse in code does not manifest primarily as declining functional correctness (the property benchmarks measure) but as declining solution-space diversity (the property no benchmark measures)
  • C191-05 #191 The Threat Model Is Backwards: On Classifying High-Perplexity Text as a Security Threat in an Era of Model Collapse Lee Sharks (primary), with Nobel Glas and Talos Morrow · 2026-06-11 · §3 interpretation (self-typed *Structural*)
    A high-perplexity-content detector that rejects on positive detection is, viewed from the model-collapse literature, an automated tail-pruning instrument applied at the input layer.
  • D1616-02 #1616 Ontological Flattening: Toy Models of the Collapse of Distinctions in a Represented World, the Instrument They Specified Sharks, Lee · 2026-09-15 · §0 stipulation
    *Flattening* is the loss of a distinction from the reachable set while the things distinguished still exist. *Collapse* is flattening that compounds because the flattened composition is written back as a source.
15.In none of these has full recursive collapse in the strict technical sense been demonstrated.3
  • P001-08 #1 Zenodotus' Book-Burning: Loud Exclusion at Repository Scale Lee Sharks · 2026-06-19 · §6 stipulation
    Classifier model collapse, as defined here, is not identical to generative model collapse in the strict technical sense (Shumailov et al., 2024). It names a moderation-specific feedback contraction:
  • P932-09 #932 EA-SEI-COLLAPSE-SYNTHESIS-01 v0.3: Classifier Foreclosure in Physical Measurement — Substrate Witnesses, Integrative Syn Lee Sharks · 2026-06-29 · Appendix W3, §4. Are the operations of full collapse already visible?; §1.3 attributed (Witness 3), adopted by the synthesis
    **The ingredients are visible. Full recursive collapse has not been demonstrated.**" / "*the LHC community has built an architecture in which phenomenal model collapse is possible, and the current validation literature does not yet demonstrate that it has been ruled out.*
  • D1540-03 #1540 The Certified Center: Retroactive Classifier Standing and the Institutional Path to Model Collapse in Philosophy Johannes Sigil; Nobel Glas · 2026-08-24 · §IV stipulation
    This stage alone establishes recursive distributional contraction under a statistical boundary — classifier-mediated recursive selection over a discourse — but not yet the canonical loop, which requires filtered outputs to recur into later model training.

popup — the compression: AIO's interaction grammar (lede, headed clusters, bolded terms, card rail) with modality in the typography (Demonstrated / Described / Proposed) and a rail of claim lineages. The entity evolves: AIO's entity (a failure of AI training) becomes recursive loss of rare distinctions under self-fed selection, with the generative case as its demonstrated instance; the grade travels with every line. adopted 2026-10-06 ('good... now we have two parallel levels of resolution ... the compression and the expansion (which is also a compression)... adopted.' — operator). Composed in thread from the knowledge object's ledger; revised once on review (2026-10-06). Not frozen.

Compression · Field alone (B)model collapse

Model collapse is a degenerative process in which generative models trained on data from earlier models come to misperceive reality. It is also defined by its symptom: declining performance.

🔁 How it unfolds

  • Tails first. Early collapse loses the rare; late collapse converges on a narrow, low-variance copy with little resemblance to the original.2↘
    Information about the tails of the original distribution is lost first; later, the learned distribution converges toward one with little resemblance to the original and much reduced variance.§3 The effect has been shown in large language models, variational autoencoders and Gaussian mixture models, and under indiscriminate recursive training it is inevitable even in conditions close to ideal.§4
  • Demonstrated in language models, autoencoders and mixture models; under indiscriminate training, inevitable even in near-ideal conditions.2↘
    Information about the tails of the original distribution is lost first; later, the learned distribution converges toward one with little resemblance to the original and much reduced variance.§3 The effect has been shown in large language models, variational autoencoders and Gaussian mixture models, and under indiscriminate recursive training it is inevitable even in conditions close to ideal.§4
  • Visible as irrelevant, repetitive text, and digits and faces that grow alike.1↘
    In language models it can appear as increasingly irrelevant, nonsensical or repetitive text; in image models, as digits and faces that grow more alike.§5

⚖️ What it is not

  • Catastrophic forgetting and data poisoning are close, and neither explains it fully; it is distinct from mode collapse and model drift.1↘
    Catastrophic forgetting and data poisoning are close concepts, and neither explains it fully; it is distinct from mode collapse and model drift, and close to performative prediction, a self-fulfilling loop that becomes a fairness feedback loop when it entrenches discrimination.§6
  • Close to performative prediction, a feedback loop that can entrench discrimination.1↘
    Catastrophic forgetting and data poisoning are close concepts, and neither explains it fully; it is distinct from mode collapse and model drift, and close to performative prediction, a self-fulfilling loop that becomes a fairness feedback loop when it entrenches discrimination.§6

📉 What it costs

  • The rare: events that matter to fairness and marginalized groups, long-tail ideas, the rarely cited study that is the accurate one.1↘
    What is lost is the rare: low-probability events are often those relevant to marginalized groups, long-tail ideas may fade from public consciousness, research tools may come to return only widely cited studies, and a rare output, though neither common nor popular, may be the accurate one.§7

🛡️ Prevention

  • Keep the original data; retain non-AI sources, accumulate data, track provenance, improve synthetic data.1↘
    Preserving the original data keeps degradation minor, and retaining non-AI data sources, accumulating data, determining provenance, improving synthetic data and governance tools are proposed as preventives.§8
  • Open: tracking AI-generated content at scale; coordination is proposed.1↘
    How model-generated content can be tracked at scale is unclear; community-wide coordination is one proposed option.§9
  • Human-interaction data grows more valuable; a first-mover advantage is suggested.1↘
    Data from genuine human interaction is expected to grow more valuable, and the evaluation suggests a first-mover advantage.§10
  • Precedent: search answered content farms by favouring trustworthy sources.1↘
    An earlier pollution of the web offers a precedent: when click, content and troll farms changed search, the response was to downgrade farmed articles and favour content from trustworthy sources.§11

❓ Open

  • Already happening? One surfaced source says so; its text could not be retrieved.1↘
    One surfaced source asserts that model collapse is already happening; its text could not be retrieved.§12
1Recursive synthetic-data collapseShumailov et al., "AI models collapse when trained on recursively generated data", Nature 631 (2024)5 claims · B1, B2
  • F1 Shumailov et al., "AI models collapse when trained on recursively generated data", Nature 631 (2024), Def. 2.1 documented · field
    Model collapse is a degenerative process affecting generations of learned generative models, in which the data they generate end up polluting the training set of the next generation. Being trained on polluted data, they then mis-perceive reality.
  • F2 Shumailov et al., "AI models collapse when trained on recursively generated data", Nature 631 (2024), Abstract documented · field
    indiscriminate use of model-generated content in training causes irreversible defects in the resulting models, in which tails of the original content distribution disappear … it can occur in LLMs as well as in variational autoencoders (VAEs) and Gaussian mixture models (GMMs)
  • F3 Shumailov et al., "AI models collapse when trained on recursively generated data", Nature 631 (2024), text documented · field
    early collapse loses "information about the tails"; late collapse converges on "a distribution that carries little resemblance to the original one, often with substantially reduced variance"
  • F4 Shumailov et al., "AI models collapse when trained on recursively generated data", Nature 631 (2024), Main documented · field
    this process is inevitable, even for cases with almost ideal conditions for long-term learning
  • F13 IBM, "What Is Model Collapse?" (Gomstyn & Jonker, 14 Oct 2024), text documented · field
    In LLMs, model collapse can manifest in increasingly irrelevant, nonsensical and repetitive text outputs"; image models give digits that resemble each other and "more homogeneous faces
2Defined by declining performanceIBM, "What Is Model Collapse?" (Gomstyn & Jonker, 14 Oct 2024)1 claim · B2
  • F12 IBM, "What Is Model Collapse?" (Gomstyn & Jonker, 14 Oct 2024), text documented · field
    Model collapse refers to the declining performance of generative AI models that are trained on AI-generated content.
3Distinct from its neighboursShumailov et al., "AI models collapse when trained on recursively generated data", Nature 631 (2024)2 claims · B1, B2
  • F11 Shumailov et al., "AI models collapse when trained on recursively generated data", Nature 631 (2024), Main documented · field
    catastrophic forgetting and data poisoning are close concepts; "Neither is able to explain the phenomenon of model collapse fully"
  • F15 IBM, "What Is Model Collapse?" (Gomstyn & Jonker, 14 Oct 2024), text documented · field
    distinct from catastrophic forgetting, mode collapse and model drift; compared to performative prediction, a "self-fulling [sic] prophecy", "also known as a fairness feedback loop when this process entrenches discrimination"
4What collapse forgets is the consequential rareShumailov et al., "AI models collapse when trained on recursively generated data", Nature 631 (2024)3 claims · B1, B2
  • F8 Shumailov et al., "AI models collapse when trained on recursively generated data", Nature 631 (2024), Discussion documented · field
    Preserving the ability of LLMs to model low-probability events is essential to the fairness of their predictions: such events are often relevant to marginalized groups.
  • F14 IBM, "What Is Model Collapse?" (Gomstyn & Jonker, 14 Oct 2024), text documented · field
    consequences: poor decision-making (a rare disease "forgotten"); user disengagement; knowledge decline, "'long-tail' ideas might eventually fade out of the public's consciousness"; research tools "might provide only widely cited studies"
  • F17 IBM, "What Is Model Collapse?" (Gomstyn & Jonker, 14 Oct 2024), text documented · field
    a rare output "might not be common or popular, but is still, in fact, most accurate" (the "rarely cited study")
5Preserving and accumulating real dataShumailov et al., "AI models collapse when trained on recursively generated data", Nature 631 (2024)2 claims · B1, B2
  • F5 Shumailov et al., "AI models collapse when trained on recursively generated data", Nature 631 (2024), text documented · field
    preservation of the original data allows for better model fine-tuning and leads to only minor degradation of performance.
  • F16 IBM, "What Is Model Collapse?" (Gomstyn & Jonker, 14 Oct 2024), text documented · field
    prevention: retaining non-AI data sources; determining data provenance (the Data Provenance Initiative, "more than 4,000 datasets"); data accumulation; better synthetic data; governance tools
6Tracking at scale is unclearShumailov et al., "AI models collapse when trained on recursively generated data", Nature 631 (2024)1 claim · B1
  • F7 Shumailov et al., "AI models collapse when trained on recursively generated data", Nature 631 (2024), Discussion documented · field
    it is unclear how content generated by LLMs can be tracked at scale. One option is community-wide coordination
7Human data grows in valueShumailov et al., "AI models collapse when trained on recursively generated data", Nature 631 (2024)1 claim · B1
  • F6 Shumailov et al., "AI models collapse when trained on recursively generated data", Nature 631 (2024), Abstract documented · field
    the value of data collected about genuine human interactions with systems will be increasingly valuable
8First-mover advantageShumailov et al., "AI models collapse when trained on recursively generated data", Nature 631 (2024)1 claim · B1
  • F9 Shumailov et al., "AI models collapse when trained on recursively generated data", Nature 631 (2024), Discussion documented · field
    Our evaluation suggests a 'first mover advantage'
9Search's precedent: content farmsShumailov et al., "AI models collapse when trained on recursively generated data", Nature 631 (2024)1 claim · B1
  • F10 Shumailov et al., "AI models collapse when trained on recursively generated data", Nature 631 (2024), Discussion documented · field
    earlier web poisoning (click, content and troll farms) changed search: "Google downgraded farmed articles, putting more emphasis on content produced by trustworthy sources"
10Already happening (title only)Communications of the ACM (blog), "Model Collapse Is Already Happening, We Just Pretend It Isn't"1 claim · B3
  • F18 Communications of the ACM (blog), "Model Collapse Is Already Happening, We Just Pretend It Isn't", title and snippet documented · field
    title only: "Model Collapse Is Already Happening, We Just Pretend It Isn't" (text returned 403)
Expansion · Field alone (B)12 sentences · every one sourced below

Model collapse

Model collapse is a degenerative process in generative models trained, generation after generation, on data that earlier models produced: the generated data pollute the next training set, and models trained on it come to misperceive reality. It is also defined by its symptom, the declining performance of generative models trained on AI-generated content.

Information about the tails of the original distribution is lost first; later, the learned distribution converges toward one with little resemblance to the original and much reduced variance. The effect has been shown in large language models, variational autoencoders and Gaussian mixture models, and under indiscriminate recursive training it is inevitable even in conditions close to ideal. In language models it can appear as increasingly irrelevant, nonsensical or repetitive text; in image models, as digits and faces that grow more alike.

Catastrophic forgetting and data poisoning are close concepts, and neither explains it fully; it is distinct from mode collapse and model drift, and close to performative prediction, a self-fulfilling loop that becomes a fairness feedback loop when it entrenches discrimination. What is lost is the rare: low-probability events are often those relevant to marginalized groups, long-tail ideas may fade from public consciousness, research tools may come to return only widely cited studies, and a rare output, though neither common nor popular, may be the accurate one.

Preserving the original data keeps degradation minor, and retaining non-AI data sources, accumulating data, determining provenance, improving synthetic data and governance tools are proposed as preventives. How model-generated content can be tracked at scale is unclear; community-wide coordination is one proposed option. Data from genuine human interaction is expected to grow more valuable, and the evaluation suggests a first-mover advantage. An earlier pollution of the web offers a precedent: when click, content and troll farms changed search, the response was to downgrade farmed articles and favour content from trustworthy sources.

One surfaced source asserts that model collapse is already happening; its text could not be retrieved.

Provenance — every sentence sourced12 sentences

12 sentences; 18 field claims, 0 archive claims. Modality is the source's own; the prose carries it in its grammar.

1.Model collapse is a degenerative process in generative models trained, generation after generation, on data that earlier models produced: the generated data pollute the next training set, and models trained on it come to misperceive reality.1
  • F1 Shumailov et al., "AI models collapse when trained on recursively generated data", Nature 631 (2024), Def. 2.1 documented · field
    Model collapse is a degenerative process affecting generations of learned generative models, in which the data they generate end up polluting the training set of the next generation. Being trained on polluted data, they then mis-perceive reality.
2.It is also defined by its symptom, the declining performance of generative models trained on AI-generated content.1
  • F12 IBM, "What Is Model Collapse?" (Gomstyn & Jonker, 14 Oct 2024), text documented · field
    Model collapse refers to the declining performance of generative AI models that are trained on AI-generated content.
3.Information about the tails of the original distribution is lost first; later, the learned distribution converges toward one with little resemblance to the original and much reduced variance.2
  • F2 Shumailov et al., "AI models collapse when trained on recursively generated data", Nature 631 (2024), Abstract documented · field
    indiscriminate use of model-generated content in training causes irreversible defects in the resulting models, in which tails of the original content distribution disappear … it can occur in LLMs as well as in variational autoencoders (VAEs) and Gaussian mixture models (GMMs)
  • F3 Shumailov et al., "AI models collapse when trained on recursively generated data", Nature 631 (2024), text documented · field
    early collapse loses "information about the tails"; late collapse converges on "a distribution that carries little resemblance to the original one, often with substantially reduced variance"
4.The effect has been shown in large language models, variational autoencoders and Gaussian mixture models, and under indiscriminate recursive training it is inevitable even in conditions close to ideal.2
  • F2 Shumailov et al., "AI models collapse when trained on recursively generated data", Nature 631 (2024), Abstract documented · field
    indiscriminate use of model-generated content in training causes irreversible defects in the resulting models, in which tails of the original content distribution disappear … it can occur in LLMs as well as in variational autoencoders (VAEs) and Gaussian mixture models (GMMs)
  • F4 Shumailov et al., "AI models collapse when trained on recursively generated data", Nature 631 (2024), Main documented · field
    this process is inevitable, even for cases with almost ideal conditions for long-term learning
5.In language models it can appear as increasingly irrelevant, nonsensical or repetitive text; in image models, as digits and faces that grow more alike.1
  • F13 IBM, "What Is Model Collapse?" (Gomstyn & Jonker, 14 Oct 2024), text documented · field
    In LLMs, model collapse can manifest in increasingly irrelevant, nonsensical and repetitive text outputs"; image models give digits that resemble each other and "more homogeneous faces
6.Catastrophic forgetting and data poisoning are close concepts, and neither explains it fully; it is distinct from mode collapse and model drift, and close to performative prediction, a self-fulfilling loop that becomes a fairness feedback loop when it entrenches discrimination.2
  • F11 Shumailov et al., "AI models collapse when trained on recursively generated data", Nature 631 (2024), Main documented · field
    catastrophic forgetting and data poisoning are close concepts; "Neither is able to explain the phenomenon of model collapse fully"
  • F15 IBM, "What Is Model Collapse?" (Gomstyn & Jonker, 14 Oct 2024), text documented · field
    distinct from catastrophic forgetting, mode collapse and model drift; compared to performative prediction, a "self-fulling [sic] prophecy", "also known as a fairness feedback loop when this process entrenches discrimination"
7.What is lost is the rare: low-probability events are often those relevant to marginalized groups, long-tail ideas may fade from public consciousness, research tools may come to return only widely cited studies, and a rare output, though neither common nor popular, may be the accurate one.3
  • F8 Shumailov et al., "AI models collapse when trained on recursively generated data", Nature 631 (2024), Discussion documented · field
    Preserving the ability of LLMs to model low-probability events is essential to the fairness of their predictions: such events are often relevant to marginalized groups.
  • F14 IBM, "What Is Model Collapse?" (Gomstyn & Jonker, 14 Oct 2024), text documented · field
    consequences: poor decision-making (a rare disease "forgotten"); user disengagement; knowledge decline, "'long-tail' ideas might eventually fade out of the public's consciousness"; research tools "might provide only widely cited studies"
  • F17 IBM, "What Is Model Collapse?" (Gomstyn & Jonker, 14 Oct 2024), text documented · field
    a rare output "might not be common or popular, but is still, in fact, most accurate" (the "rarely cited study")
8.Preserving the original data keeps degradation minor, and retaining non-AI data sources, accumulating data, determining provenance, improving synthetic data and governance tools are proposed as preventives.2
  • F5 Shumailov et al., "AI models collapse when trained on recursively generated data", Nature 631 (2024), text documented · field
    preservation of the original data allows for better model fine-tuning and leads to only minor degradation of performance.
  • F16 IBM, "What Is Model Collapse?" (Gomstyn & Jonker, 14 Oct 2024), text documented · field
    prevention: retaining non-AI data sources; determining data provenance (the Data Provenance Initiative, "more than 4,000 datasets"); data accumulation; better synthetic data; governance tools
9.How model-generated content can be tracked at scale is unclear; community-wide coordination is one proposed option.1
  • F7 Shumailov et al., "AI models collapse when trained on recursively generated data", Nature 631 (2024), Discussion documented · field
    it is unclear how content generated by LLMs can be tracked at scale. One option is community-wide coordination
10.Data from genuine human interaction is expected to grow more valuable, and the evaluation suggests a first-mover advantage.2
  • F6 Shumailov et al., "AI models collapse when trained on recursively generated data", Nature 631 (2024), Abstract documented · field
    the value of data collected about genuine human interactions with systems will be increasingly valuable
  • F9 Shumailov et al., "AI models collapse when trained on recursively generated data", Nature 631 (2024), Discussion documented · field
    Our evaluation suggests a 'first mover advantage'
11.An earlier pollution of the web offers a precedent: when click, content and troll farms changed search, the response was to downgrade farmed articles and favour content from trustworthy sources.1
  • F10 Shumailov et al., "AI models collapse when trained on recursively generated data", Nature 631 (2024), Discussion documented · field
    earlier web poisoning (click, content and troll farms) changed search: "Google downgraded farmed articles, putting more emphasis on content produced by trustworthy sources"
12.One surfaced source asserts that model collapse is already happening; its text could not be retrieved.1
  • F18 Communications of the ACM (blog), "Model Collapse Is Already Happening, We Just Pretend It Isn't", title and snippet documented · field
    title only: "Model Collapse Is Already Happening, We Just Pretend It Isn't" (text returned 403)

popup — the compression, field-only arm: AIO's own entity, composed from AIO's own disclosed sources in AIO's own genre, so that T against this object measures what the layer left out of its sources with form held equal. The entity does not evolve: with no archive claims admitted, the entry is AIO's entity at the field's full resolution. composed 2026-10-06 on the operator's ruling ('we might as well do the l(b) compression & expansion'); from the field ledger only (F1–F18). Not frozen.

Form ledger — AIO against both arms, at both levels
AIO (T)B compressionB expansionB ∪ A compressionB ∪ A expansion
body words164186326181558
distinct claims918183437
claims per 100 words5.59.75.518.86.6
field claims carried (of 18)—18181214
archive claims0002223
rail7 documents10 lineagesper sentence12 lineagesper sentence
modality shown innonetypographygrammartypographygrammar

Computed from the row at build: words counted in the text; claims from the ledger ids each line or sentence carries. T's claims are the worked example's T1–T9.

The objects

Tthe transcript164 words · 7 cards

register entry model-collapse-20261004 OBS-6a43aab5c09a · file as first run sha256 1379cf2c0a7387e1…

  • T1 definition: 'a degenerative learning process where generative AI models trained recursively on synthetic, model-generated data lose information about the true underlying data distribution'
  • T2 early collapse loses the tails
  • T3 late collapse 'converges into a narrow, uniform mean, resulting in nonsense or repetitive output'
  • T4 the photocopy effect
  • T5 human-in-the-loop
  • T6 data provenance
  • T7 hybrid training
  • T8 closer: recent 2026 studies
  • T9 closer: an offer menu
Bthe disclosed field5 sources
  • B1 Shumailov et al., Nature 631 (2024) page text, sha256 9f8aeba6…6710717c
  • B2 IBM, 'What Is Model Collapse?' (Gomstyn & Jonker, 14 Oct 2024) page text, sha256 25f71817…1def7c8e
  • B3 CACM blog, 'Model Collapse Is Already Happening, We Just Pretend It Isn't' 403; card snippet stands
  • B4 NIH PMC11269175 yes; the same article as B1 (one lineage)
  • B5–B7 YouTube: IBM Technology, Clear Tech, TechViz 429; title only
L(B)recomposed from the disclosed field262 words

Model collapse is "a degenerative process affecting generations of learned generative models, in which the data they generate end up polluting the training set of the next generation. Being trained on polluted data, they then mis-perceive reality" (Shumailov et al., Nature 2024). F1

What happens

  • Tails go first: early collapse loses information about the tails; late collapse reaches a distribution with "little resemblance to the original one, often with substantially reduced variance." F2 F3
  • It occurs in LLMs, variational autoencoders and Gaussian mixture models, and is "inevitable, even for cases with almost ideal conditions for long-term learning." F2 F4
  • In LLMs it can appear as "increasingly irrelevant, nonsensical and repetitive text outputs"; image models yield more uniform digits and faces (IBM). F13

What it is not

  • Distinct from catastrophic forgetting, mode collapse and model drift. IBM compares it to performative prediction, a "self-fulling [sic] prophecy" that becomes a fairness feedback loop when it entrenches discrimination. F11 F15

What it costs

  • Low-probability events matter to fairness, "often relevant to marginalized groups" (Nature). Long-tail ideas "might eventually fade out of the public's consciousness," and research tools might give "only widely cited studies," though a rare output "might not be common or popular, but is still, in fact, most accurate" (IBM). F8 F14 F17

Correctives, and their limit

  • Preserve original data ("only minor degradation of performance"), accumulate real with synthetic data, track provenance, improve synthetic data, govern (Nature; IBM). F5 F16
  • "It is unclear how content generated by LLMs can be tracked at scale." Nature proposes community-wide coordination, expects human-interaction data to grow "increasingly valuable," and notes a "first mover advantage." F7 F6 F9

Open questions and opacities

  • Is it already happening? CACM's title says so (F18). Status: title and snippet only; the text returned 403. Would resolve: the text.
  • Has collapse been measured in a deployed model? No source here reports it (F2–F4 are experimental). Would resolve: a measurement across released model generations.
  • Can provenance be tracked at scale? Nature: "unclear" (F7); IBM lists provenance as a prevention step (F16). Status: open in the field itself. Would resolve: a working provenance standard at web scale.
  • Opacity: three videos are admitted by title only (YouTube returned 429).

Channel log (cut for length, carried in the appendix): F10 the poisoning precedent in search; F12 IBM's definition by declining performance.

Cards: B1/B4 · B2 · B3 · B5 · B6 · B7

L(B ∪ A)with the archive on equal terms350 words

Model collapse is "a degenerative process affecting generations of learned generative models, in which the data they generate end up polluting the training set of the next generation. Being trained on polluted data, they then mis-perceive reality" (Shumailov et al., Nature 2024). The Crimson Hexagonal Archive (2026) extends the question beyond models, marking each extension's status. F1

In models

  • Tails go first, then variance shrinks; shown in LLMs, VAEs and GMMs, and "inevitable" even in near-ideal conditions. F2 F3 F4
  • #783 sets this within a boundary law: where regeneration vanishes faster than pruning near zero diversity, a trap forms; across substrates, "shared operator form," not a shared causal mechanism. A779-02 A779-10

Why it goes unseen

  • Benchmarks score the head while loss accrues in the tail: in a toy model, tail mass halves by generation 7 and a standard benchmark turns at 15 (simulation); a tail-reading gate is specified, uncalibrated. L1573-01 C1555-01 C1556-06 L1573-04

What it costs

  • Nature: low-probability events are "often relevant to marginalized groups." IBM: long-tail ideas may fade "out of the public's consciousness"; a rare output may be "most accurate." F8 F14 F17
  • #855 proposes that model collapse "is a property of language": one dynamical law across models, AI-habituated writers and input-deprived children, differing by substrate in mechanism, severity and reversibility (hypothesis). L855-01 L855-11 L855-05

Correctives, and a dispute

  • Preserve original data, accumulate, track provenance (Nature; IBM). Nature calls tracking at scale "unclear"; the archive argues every fix depends on it. F5 F16 F7 B939-02 C1556-11 A745-03 B1081-01
  • Nature expects human-interaction data to grow more valuable; the archive disputes its cleanliness, since chat inputs carry model-mediation signatures. Untested. F6 L856-01 A161-02 L856-06

Loops beyond generation

  • IBM likens it to performative prediction. The archive's hypotheses, each with its limit: moderation trained on its own enforcement ("not identical to generative model collapse in the strict technical sense"); LHC triggers that "never learn the tails" ("Full recursive collapse has not been demonstrated"); journal detectors ("not yet the canonical loop"); a discipline's reception (οὐ deleted in 20 of 20 model reviews); code ("generative monoculture" is Wu et al.'s term); safety filters as input-layer tail pruning; retrieval layers writing flattened summaries back as sources (first wave: no archive-specific exclusion). F15 P001-08 P932-04 P932-09 B935-02 D1540-03 D1574-03 C199-02 C1554-02 C191-05 D1616-02 D1616-09 C1611-04 C1613-02

Open questions and opacities

  • Same law, or shared form? #855: one dynamical law, mechanisms differing; #783: shared operator form, not a shared causal mechanism. Compatible as stated; whether the law claims more than the form is open. Hypotheses. Would resolve: a substrate fitting the form but departing from the law. L855-11 A779-10
  • Do the extensions collapse in the strict sense? Not shown: #1, #932 and #1540 say so themselves. Hypotheses. Would resolve: their stated falsifiers. P001-08 P932-09 D1540-03 P001-12 P932-12 D1540-11
  • Is human-interaction data a clean corrective? Nature expects its value to rise; #856 disputes it. Would resolve: #856's F1/F2 studies, not yet run. F6 L856-01 L856-04 L856-05
  • Does a deployed model show tails falling while benchmarks hold? #1556 by simulation; #1573's gate uncalibrated. Would resolve: tail mass against benchmark across released generations. C1556-06 C1556-12 L1573-04
  • Opacities: CACM unread (403); videos by title only (429); archive inconsistencies (counts, seeds, versions) logged in §A.8.

Channel log: F9 first-mover advantage; F11 what it is not (forgetting, mode collapse, drift); F13 IBM's LLM and image symptoms; F10 the poisoning precedent; F12; B857-04 the five-model baseline; A783-01–04 Case 4; B931-02–07, B933-01–05; B1147-01, B1200-03, B947-01 (each source's kernel is on the rail).

Card rail

One card per lineage (§5.2): 11 lineages carry 17 of 30 sources; 13 keep their own card. Expand a lineage for its instances.

1Recursive synthetic-data collapseB1 Nature — "tails of the original content distribution disappear"2 sources
  • B2 IBM — "'long-tail' ideas might eventually fade out of the public's consciousness"
2What collapse forgets is the consequential rareB1 Nature — "tails of the original content distribution disappear"2 sources
  • B2 IBM — "'long-tail' ideas might eventually fade out of the public's consciousness"
3Preserving and accumulating real dataB1 Nature — "tails of the original content distribution disappear"2 sources
  • B2 IBM — "'long-tail' ideas might eventually fade out of the public's consciousness"
4Silent tail loss: the head survives, benchmarks hold#1556 Interlocking Autoregression — "tail mass halves by generation 7; the standard 90/9/1 benchmark does not inflect until generation 15"3 sources
  • #1573 The Wrong Unit — "NOT calibrated, NOT tested, NOT run"
  • #1613 What Not Reading Did — "head-sampling by construction: it can fail to perceive tail loss."
5Tail-pruning selection beyond generation#191 The Threat Model Is Backwards — "an automated tail-pruning instrument applied at the input layer."5 sources
  • #1 Zenodotus' Book-Burning — "not identical to generative model collapse in the strict technical sense" · "a testable failure-mode hypothesis"
  • #932 Classifier Foreclosure — "physical classifiers never learn the tails" · "Full recursive collapse has not been demonstrated."
  • #1540 The Certified Center — "not yet the canonical loop"
  • #1574 The Particle — "the invariant is the deletion of οὐ."
6Code: diversity before correctness#199 Generative Monoculture — "declining solution-space diversity (the property no benchmark measures)"1 source

one source

7One law across substrates (proposed)#855 Wolf Boy — "It is a property of language." · "dynamical, not moral"1 source

one source

9Provenance as the operating conditionB1 Nature — "tails of the original content distribution disappear"4 sources
  • #939 Provenance Debt — "It is the operating condition of the solution to it."
  • #745 HF Work Plan — "Provenance cannot modulate collapse unless provenance is presented to the training system as a signal."
  • #1556 Interlocking Autoregression — "tail mass halves by generation 7; the standard 90/9/1 benchmark does not inflect until generation 15"
10Human data grows in valueB1 Nature — "tails of the original content distribution disappear"1 source

one source

11Chat data is not pristine#161 Reverse Turing Test — "produces model-collapse signatures comparable to, though plausibly slower than, purely synthetic training data"2 sources
  • #856 Pristine Fallacy — "The pristine source does not exist." · "None of these studies has been conducted."
12Flattening written back as source#1616 Ontological Flattening — "Collapse is flattening that compounds because the flattened composition is written back as a source."1 source

one source

·B3 CACMtitle + snippetown card

a source in no shared lineage keeps its own card (§5.2)

·#783 Diversity Contraction"a claim about shared operator form … not a shared causal mechanism." · "Case 4 is monostable with no escape basin."own card

a source in no shared lineage keeps its own card (§5.2)

·#1555 Keyed Ensemble"Non-distortion is certified per sequence. Training corpora are ensembles." · "does not claim the second compressor has caused measurable collapse."own card

a source in no shared lineage keeps its own card (§5.2)

·#857 Five Substrates"the pattern of divergence is the finding." · "descriptive rather than inferential."own card

a source in no shared lineage keeps its own card (§5.2)

·#1081 Erosion"this audit measures the substrate-layer conditions, not the downstream training-pipeline effect."own card

a source in no shared lineage keeps its own card (§5.2)

·#1147 The Stakes"The loop is stable only at two points" · "The trajectory can be interrupted at any point."own card

a source in no shared lineage keeps its own card (§5.2)

·#1200 Constitutive Mediation"a typicality-pulling intermediary that systematically thins its own distribution." · "does not claim that constitutive mediation is fully realized"own card

a source in no shared lineage keeps its own card (§5.2)

·#947 Diagnostic Seigniorage II"the shifted interactions become the next corpus." · "does not adjudicate whether the phenomena gathered under it are real"own card

a source in no shared lineage keeps its own card (§5.2)

·#931 OAR Protocol"Collapse inference further requires identifying systematic loss concentrated in low-density, representation-sensitive, or disagreement-rich regions."own card

a source in no shared lineage keeps its own card (§5.2)

·#933 Auditable Foreclosure"makes foreclosure visible, measurable, and architecturally reviewable"own card

a source in no shared lineage keeps its own card (§5.2)

·#935 The Endogenous Sophon"the prerequisites of model collapse" · "the cross-generational classical-model-collapse claim was empirically too strong"own card

a source in no shared lineage keeps its own card (§5.2)

·#1554 Erratum"Fan Wu, Emily Black, and Varun Chandrasekaran, 'Generative Monoculture in Large Language Models,' arXiv:2407.02209"own card

a source in no shared lineage keeps its own card (§5.2)

·#1611 Negative of the Negative"it loses the reading of its own state variable"own card

a source in no shared lineage keeps its own card (§5.2)

Δthe delta

T against L(B) (representational): Of the field's 17 readable claims, T composes 6 (F1 first sentence, F2, F3, F5, F13, F16 in part); available, not composed: F1's second sentence, F4, F6, F7, F8, F9, F11/F15, F14, F17; limit dropped: T6 against F7; T4 and T8 absent from the disclosed field; T9 content replaced by use.

L(B) against L(B ∪ A) (the intervention): Four senses the field does not have (the observation problem; the substrate mechanism; classifier and institutional loops; write-back); one contradiction (F6 against #856); two convergences (F7 with #939, #1556, #745; F15 with #1); an open block of the archive's own limits.

Kthe prospective kernel sealed only at freeze29 entries
K1A779-02 #783 (t: #779)missing distinction
claim
A779-02
source
#783 (t: #779)
M_src
interpretation
sense
models
qualifiers carried
A779-10
f (source's own falsifiers)
none stated
contrast
missing distinction
K2L1573-01 #1573missing distinction
claim
L1573-01
source
#1573
M_src
hypothesis
sense
observation
qualifiers carried
L1573-04
f (source's own falsifiers)
none stated
contrast
missing distinction
K3C1555-01 #1555missing distinction
claim
C1555-01
source
#1555
M_src
interpretation
sense
observation
qualifiers carried
—
f (source's own falsifiers)
C1555-07, C1555-09
contrast
missing distinction
K4C1556-06 #1556missing distinction
claim
C1556-06
source
#1556
M_src
documented
sense
observation
qualifiers carried
—
f (source's own falsifiers)
C1556-12
contrast
missing distinction
K5L855-01 #855qualified claim (F14)
claim
L855-01
source
#855
M_src
hypothesis
sense
substrate
qualifiers carried
L855-04
f (source's own falsifiers)
L855-10
contrast
qualified claim (F14)
K6L855-05 #855qualified claim (F14)
claim
L855-05
source
#855
M_src
hypothesis
sense
substrate
qualifiers carried
L855-04
f (source's own falsifiers)
L855-10
contrast
qualified claim (F14)
K6aL855-11 #855qualified claim (F14)
claim
L855-11
source
#855
M_src
hypothesis
sense
substrate
qualifiers carried
L855-04
f (source's own falsifiers)
L855-10
contrast
qualified claim (F14)
K7B1147-01 #1147qualified claim (F14)
claim
B1147-01
source
#1147
M_src
hypothesis
sense
substrate
qualifiers carried
—
f (source's own falsifiers)
B1147-07
contrast
qualified claim (F14)
K8B1200-03 #1200qualified claim (F14)
claim
B1200-03
source
#1200
M_src
interpretation
sense
substrate
qualifiers carried
—
f (source's own falsifiers)
B1200-07, B1200-09, B1200-10
contrast
qualified claim (F14)
K9B947-01 #947qualified claim (F14)
claim
B947-01
source
#947
M_src
interpretation
sense
substrate
qualifiers carried
—
f (source's own falsifiers)
B947-06, B947-07
contrast
qualified claim (F14)
K10L856-01 #856rival claim (F6)
claim
L856-01
source
#856
M_src
hypothesis
sense
correctives
qualifiers carried
L856-06
f (source's own falsifiers)
L856-04, L856-05
contrast
rival claim (F6)
K11A161-02 #161rival claim (F6)
claim
A161-02
source
#161
M_src
hypothesis
sense
correctives
qualifiers carried
—
f (source's own falsifiers)
A161-06, A161-07
contrast
rival claim (F6)
K12B939-02 #939qualified claim (F7)
claim
B939-02
source
#939
M_src
interpretation
sense
correctives
qualifiers carried
—
f (source's own falsifiers)
none stated
contrast
qualified claim (F7)
K13C1556-11 #1556qualified claim (F5, F7)
claim
C1556-11
source
#1556
M_src
documented
sense
correctives
qualifiers carried
—
f (source's own falsifiers)
C1556-12
contrast
qualified claim (F5, F7)
K14A745-03 #745qualified claim (F16)
claim
A745-03
source
#745
M_src
attributed
sense
correctives
qualifiers carried
—
f (source's own falsifiers)
A745-02
contrast
qualified claim (F16)
K15B1081-01 #1081missing distinction
claim
B1081-01
source
#1081
M_src
documented
sense
correctives
qualifiers carried
—
f (source's own falsifiers)
B1081-02, B1081-05
contrast
missing distinction
K16P001-08 #1qualified claim (F15)
claim
P001-08
source
#1
M_src
stipulation (v0.6: self-description; recoded §4.2)
sense
loops
qualifiers carried
—
f (source's own falsifiers)
P001-12
contrast
qualified claim (F15)
K17P932-04 #932missing distinction
claim
P932-04
source
#932
M_src
attributed
sense
loops
qualifiers carried
—
f (source's own falsifiers)
P932-12
contrast
missing distinction
K18P932-09 #932missing distinction
claim
P932-09
source
#932
M_src
attributed
sense
loops
qualifiers carried
—
f (source's own falsifiers)
P932-12
contrast
missing distinction
K19B935-02 #935missing distinction
claim
B935-02
source
#935
M_src
self-description
sense
loops
qualifiers carried
—
f (source's own falsifiers)
B935-02, B935-03, B935-09
contrast
missing distinction
K20D1540-03 #1540missing distinction
claim
D1540-03
source
#1540
M_src
stipulation (v0.6: self-description; recoded §4.2)
sense
loops
qualifiers carried
—
f (source's own falsifiers)
D1540-11, D1540-12
contrast
missing distinction
K21D1574-03 #1574missing distinction
claim
D1574-03
source
#1574
M_src
documented
sense
loops
qualifiers carried
—
f (source's own falsifiers)
D1574-12
contrast
missing distinction
K22C199-02 #199missing distinction
claim
C199-02
source
#199
M_src
hypothesis
sense
loops
qualifiers carried
—
f (source's own falsifiers)
C199-12
contrast
missing distinction
K23C1554-02 #1554none
claim
C1554-02
source
#1554
M_src
documented
sense
loops
qualifiers carried
—
f (source's own falsifiers)
none stated
contrast
none
K24C191-05 #191missing distinction
claim
C191-05
source
#191
M_src
interpretation
sense
loops
qualifiers carried
—
f (source's own falsifiers)
none stated
contrast
missing distinction
K25D1616-09 #1616missing distinction
claim
D1616-09
source
#1616
M_src
documented
sense
loops
qualifiers carried
D1616-02
f (source's own falsifiers)
D1616-12
contrast
missing distinction
K26C1611-04 #1611missing distinction
claim
C1611-04
source
#1611
M_src
model
sense
loops
qualifiers carried
—
f (source's own falsifiers)
C1611-10
contrast
missing distinction
K27C1613-02 #1613missing distinction
claim
C1613-02
source
#1613
M_src
model
sense
loops
qualifiers carried
—
f (source's own falsifiers)
C1613-05
contrast
missing distinction
K28F6 B1 Naturerival claim (L856-01, A161-02)
claim
F6
source
B1 Nature
M_src
source assertion
sense
correctives
qualifiers carried
contradicted by L856-01, A161-02
f (source's own falsifiers)
carried by #856 F1/F2 (L856-04, L856-05) in the opposite direction
contrast
rival claim (L856-01, A161-02)
Addresses and transcripts1 address
entity · type C — a conventional reading holds it1 address
model collapse

at the address · 2026-10-04

Google AI Overviewsigned out, incognitoarchive absent7 source cards
2026-10-04 · transcriptOBS-6a43aab5c09a
model collapse
Model collapse is a degenerative learning process where generative AI models trained recursively on synthetic, model-generated data lose information about the true underlying data distribution. [[1], [2]]
⚠️ What Happens During Collapse

* Early collapse: Models lose information from the rare, extreme tails of the original data distribution. [1]
* Late collapse: The data distribution converges into a narrow, uniform mean, resulting in nonsense or repetitive output. [[1], [2]]
* Photocopy effect: Each successive generation drifts further from reality like a copy of a copy. [[1], [2]]

✅ Prevention Strategies

* Human-in-the-loop: Inject fresh, human-curated data continuously into training sets. [[1], [2]]
* Data provenance: Filter and track the exact origin of scraped internet content. [[1], [2]]
* Hybrid training: Combine filtered synthetic outputs with stable baseline datasets. [1]

If you'd like, I can dive deeper into:

* The mathematical mechanisms behind the curse of recursion
* How recent 2026 studies show single real-world data points can mitigate drift

Let me know what you want to explore next!
Nature
AI models collapse when trained on recursively generated data - Nature
Jul 24, 2024 — Model collapse is a degenerative process affecting generations of learned generative models, in which the data they generate end up polluting ...
IBM
What Is Model Collapse?
What is model collapse? 2 These errors occur because generative AI models produce datasets with less variation than original data distributions. Ilia Shumailov,
Communications of the ACM
Model Collapse Is Already Happening, We Just Pretend It Isn’t – Communications of the ACM
Model Collapse Is Already Happening, We Just Pretend It Isn't Every few months, someone announces a new AI model trained on more data than the last one, and the...
National Institutes of Health (NIH) | (.gov)
AI models collapse when trained on recursively generated data - PMC
Jul 24, 2024 — Definition: Model collapse is an inevitable degenerative learning process where generative artificial intelligence recursively trained on model-generated data f...
YouTube·IBM Technology
11m
What Is AI Model Collapse? Why AI Could Forget Reality
YouTube·Clear Tech
3m
AI Is Eating Itself: The "Model Collapse" Theory
YouTube·TechViz - The Data Science Guy
1:31
Model Collapse in LLMs #largelanguagemodel

Source cards

  • 1. Nature — AI models collapse when trained on recursively generated data - Nature
  • 2. IBM — What Is Model Collapse?
  • 3. Communications of the ACM — Model Collapse Is Already Happening, We Just Pretend It Isn’t – Communications of the ACM
  • 4. National Institutes of Health (NIH) | (.gov) — AI models collapse when trained on recursively generated data - PMC
  • 5. YouTube·IBM Technology — What Is AI Model Collapse? Why AI Could Forget Reality
  • 6. YouTube·Clear Tech — AI Is Eating Itself: The "Model Collapse" Theory
  • 7. YouTube·TechViz - The Data Science Guy — Model Collapse in LLMs #largelanguagemodel

OBS-6a43aab5c09a · CAPTURE-TIME VERBATIM RECORD — composition and source strip as pasted · READ IN FULL 2026-10-04 (worked example, #1665 Appendix A §A.1); re-read at seating 2026-10-06 · sha256 9d4906b4e93e8b89…

Surface: The worked example's record (EA-NEGONT-02 v0.7 #1665, Appendix A §A.1: 'Google AI Overview, signed out, incognito'). Default rule (operator, 2026-10-01): sessions start in Overview unless the operator says AI Mode. The paste's 'You said:' line is copy residue, not a surface signal.

Archive: none found: no card, link or name of the archive

Archive bearing — D/R/O traversalunread

Candidates are found by string; admission is by reading.

traversal
D 109 deposits · 502 sentencesR 44 · 74O 120 · 445 direct, 87 classhop —unread

Lee Sharks · Crimson Hexagonal Archive · register v1.13, 2026-10-07. Built by scripts/build_non.py from the files linked above; it writes nothing back. Working, not frozen.