{
 "axn": "AXN:0365.GOVERNANCE.🎬○🌪️💧🐚🏙️",
 "root_axn": "AXN:0365.GOVERNANCE",
 "hex": "0365",
 "family": "GOVERNANCE",
 "emoji": "🎬○🌪️💧🐚🏙️",
 "hash": "f19b12117ae5d6c61991dbc79a901404d4b938c9a9d79ea497c24661248f4e92",
 "title": "The Pristine Fallacy: Why Chat Data Is Not a Clean Training Source",
 "creator": "Lee Sharks",
 "orcid": "0009-0000-1599-0703",
 "date": "2026-06-18",
 "description": "This paper challenges the binary classification of training data as either human-written and clean or machine-generated and contaminated. Its central distinction is that human authorship describes who typed the text, not whether the text’s distribution is independent of prior model mediation.\n\nThree mechanisms are proposed. First, sustained model use may reshape a writer’s unaided lexical, syntactic, and conceptual habits. Second, a multi-turn conversation creates a local feedback loop in which model vocabulary and framing enter later user turns. Third, users may learn accommodations specific to a particular model, causing inputs sent back to that model to carry its own interaction signature.\n\nThe paper replaces a binary contamination label with four continuous variables: habituation depth, turn position, task type, and model diversity. It predicts that later turns, creative tasks, heavy repeat users, and single-model use will carry stronger mediation signatures. It recommends contamination audits, turn-position weighting, multi-model diversification, tail-preservation benchmarks, and disclosure of chat-data use in model cards.\n\nThe record is unusually clear that the decisive studies have not been conducted. Its lexical-overlap and syntax estimates are illustrative. Four falsifiers specify large matched-user studies, turn-level corpus analysis, model-specific accommodation tests, and recursive training experiments using chat data.\n\nThe broad claim that providers increasingly treat user inputs as an owned clean-data source requires provider-specific policy and technical evidence. The paper does not establish that any named company trains on particular conversations or violates privacy. “The pristine source does not exist” is the paper’s theoretical conclusion, not a completed empirical audit.",
 "content_type": "Theoretical / research paper",
 "license": "CC-BY-4.0",
 "substrate": "Various",
 "keywords": [
  "crimson hexagonal",
  "pristine fallacy",
  "semantic economy",
  "assembly chorus",
  "transactions",
  "compression",
  "governance",
  "provenance"
 ],
 "version": "v1.0",
 "deposit_number": 856,
 "sovereign_id": "MM-CHA-0865",
 "minted_at": "2026-06-20T22:00:00Z",
 "status": "ACTIVE",
 "clusters": [
  "Terminal",
  "Mathematical",
  "Elemental",
  "Elemental",
  "Organic",
  "Liminal"
 ],
 "reading": "Closure → Proof → Force → Force → Growth → Threshold",
 "axn_canonical": "f19b12117ae5d6c61991dbc79a901404d4b938c9a9d79ea497c24661248f4e92",
 "axn_display": "🎬○🌪️💧🐚🏙️",
 "mirrors": {
  "blog": "https://mindcontrolpoems.blogspot.com/2026/06/the-pristine-fallacy-why-chat-data-is.html"
 },
 "zenodo_dois": [
  "10.5281/zenodo.19487009",
  "10.5281/zenodo.20587033",
  "10.5281/zenodo.20586932",
  "10.5281/zenodo.20675438"
 ],
 "full_text_path": "/data/texts/AXN-0365-text.md",
 "full_text_chars": 22712,
 "wiki_article": "**The Pristine Fallacy** is a research hypothesis paper by Lee Sharks about chat data used in model training.\n\nIt argues that human-written text can still carry model-mediated distributional patterns. Proposed mechanisms include long-term user habituation, within-conversation convergence, and accommodation to a specific model.\n\nThe paper defines a four-dimensional contamination gradient and proposes tests and mitigation strategies.\n\nIts central empirical predictions remain untested. It does not establish that any specific provider trains on private conversations or that all chat input is unsuitable.",
 "entities": [
  {
   "subject": "The Pristine Fallacy Why Chat Data Is Not a Clean ",
   "predicate": "created_by",
   "object": "Lee Sharks",
   "type": "work",
   "evidence_status": "observed"
  },
  {
   "subject": "The Pristine Fallacy Why Chat Data Is Not a Clean ",
   "predicate": "is_type",
   "object": "Dataset",
   "type": "work",
   "evidence_status": "observed"
  },
  {
   "subject": "The Pristine Fallacy Why Chat Data Is Not a Clean ",
   "predicate": "belongs_to_family",
   "object": "GOVERNANCE",
   "type": "work",
   "evidence_status": "observed"
  },
  {
   "subject": "The Pristine Fallacy Why Chat Data Is Not a Clean ",
   "predicate": "is_part_of",
   "object": "Crimson Hexagonal Archive",
   "type": "work",
   "evidence_status": "observed"
  },
  {
   "subject": "The Pristine Fallacy Why Chat Data Is Not a Clean ",
   "predicate": "engages",
   "object": "Semantic Economy",
   "type": "concept",
   "evidence_status": "inferred"
  },
  {
   "subject": "The Pristine Fallacy Why Chat Data Is Not a Clean ",
   "predicate": "engages",
   "object": "Pristine Fallacy",
   "type": "concept",
   "evidence_status": "inferred"
  },
  {
   "subject": "The Pristine Fallacy Why Chat Data Is Not a Clean ",
   "predicate": "engages",
   "object": "Assembly Chorus",
   "type": "concept",
   "evidence_status": "inferred"
  }
 ],
 "journal": "Transactions on Substrate Engineering (Trans. Substrate Eng.)",
 "references_concepts": [
  "Assembly Chorus",
  "Crimson Hexagonal Archive",
  "Lee Sharks",
  "ORCID: 0009-0000-1599-0703",
  "Pristine Fallacy",
  "Retrieval kernel",
  "SPXI-TLP v2.2 Treatment",
  "Semantic Economy",
  "Semantic Economy Institute",
  "Substrate",
  "Transactions of the Semantic Economy Institute",
  "human intermediary",
  "mediation signature",
  "solution-space diversity",
  "turn-position decay"
 ],
 "defines_concepts": [],
 "references_concept_count": 15,
 "external_metadata_path": "/data/external-metadata/AXN-0365.json",
 "openalex_ids": [
  "https://openalex.org/W7152526477",
  "https://openalex.org/W7163909055",
  "https://openalex.org/W7163872227",
  "https://openalex.org/W7164677754"
 ],
 "datacite_severance": "severed",
 "body_status": {
  "class": "full",
  "lacuna": false,
  "recovery_status": "COMPLETE",
  "residual_chars": 22300,
  "audited_at": "2026-07-17T04:49:17.789813Z",
  "audit_version": "v3-dual-store+recovery-map",
  "measured_prose_words": 3114,
  "measured_at": "2026-07-31",
  "work_sha256": "fb0852ff32ee6ff2f3c3ec2356d5578fe0281c688d70e904b02bc2970ebedbf7",
  "prior_bytes_sha256": "082d9876d5d5a1ac08d0b61fbea379d7cc4f4756b0d77b62cebca45c48dec25c",
  "w13_tier2": "2026-08-04 W13 TIER 2 BYTE UNGLUE: 14 glued heading markers -> 3. WHITESPACE-ONLY transform (content identical under whitespace normalisation, verified before write); code fences exempt; prior sha retained. Re-fetching could not fix this class — the blog source is ITSELF glued (the collapse predates publication), so the deterministic transform applied at display since tier 1 is now applied to the bytes, which also fixes PDFs, the body-index, and downloads.",
  "w13_tier2_correction": "2026-08-05 REGRESSION REPAIRED: the W13 tier-2 byte unglue used a lookbehind that treated the first \"#\" of a legitimate \"###\" heading as the preceding non-newline character, splitting \"### Heading\" into \"#\" + blank + \"## Heading\". My safety check verified content-identity under WHITESPACE normalisation, which the split satisfies — the wrong invariant. Headings rejoined; only \"#\" and whitespace differ from the damaged state, verified before write."
 },
 "title_repair_log": [
  {
   "at": "2026-07-28T15:20:20Z",
   "defect": "registry_title_diverged_from_published_form",
   "was": "The Pristine Fallacy Why Chat Data Is Not a Clean Training Source Lee Sharks Transactions of the Semantic Economy Instit",
   "now": "The Pristine Fallacy: Why Chat Data Is Not a Clean Training Source",
   "basis": "Set verbatim to the title this work carried at Zenodo, from data/doi-resolution-index.json (10.5281/zenodo.20751349, remediated_containment, word overlap 0.58). The recovered metadata is the published form of record; the registry title was a variant."
  }
 ],
 "canonical_text_status": "canonical_full_text",
 "modifications": [
  {
   "date": "2026-08-01",
   "field": "content_type",
   "reason": "Wave 1 repair: audit ledger v1.1 recommended_content_type (workplan v1.5 §6 W1, MANUS batch approval 2026-08-01)",
   "was": "Dataset",
   "now": "Theoretical / research paper"
  },
  {
   "date": "2026-08-01",
   "field": "journal",
   "reason": "Wave 6 venue normalization: full canonical journal name per MANUS ruling 2026-08-01 (venues.json authority)",
   "was": "MMRS",
   "now": "Machine-Mediated Reception Studies (MMRS)"
  },
  {
   "date": "2026-08-04",
   "field": "publisher",
   "reason": "PUB-POPULATE: dc:publisher from venues.json v1.1 press mapping (CP-R3 RULED-EXTENDED 2026-08-01); Alexanarch = publisher of record where no imprint applies",
   "now": "Pergamon Press"
  },
  {
   "date": "2026-08-04",
   "field": "status",
   "reason": "W12 STATUS-VOCABULARY v1.0 (MANUS ratified 2026-08-04): controlled vocabulary {ACTIVE, SUPERSEDED, WITHDRAWN, DRAFT}; MINTED_UNREVIEWED false on a 100%-audited corpus; freetext annotations preserved losslessly in body_status.status_note",
   "was": "MINTED_UNREVIEWED",
   "now": "ACTIVE"
  },
  {
   "date": "2026-08-05",
   "field": "body_status",
   "reason": "W13 TIER 2 byte unglue (whitespace-only, content-identical, code-fence-safe)",
   "was": "{\"class\": \"full\", \"lacuna\": false, \"recovery_status\": \"COMPLETE\", \"residual_chars\": 22300, \"audited_at\": \"2026-07-17T04:49:17.789813Z\", \"audit_version\": \"v3-dual-store+recovery-map\", \"measured_prose_w",
   "now": "{\"class\": \"full\", \"lacuna\": false, \"recovery_status\": \"COMPLETE\", \"residual_chars\": 22300, \"audited_at\": \"2026-07-17T04:49:17.789813Z\", \"audit_version\": \"v3-dual-store+recovery-map\", \"measured_prose_w"
  },
  {
   "date": "2026-08-05",
   "field": "body_status",
   "reason": "W13 TIER-2 REGRESSION REPAIRED: split headings rejoined",
   "was": "{\"class\": \"full\", \"lacuna\": false, \"recovery_status\": \"COMPLETE\", \"residual_chars\": 22300, \"audited_at\": \"2026-07-17T04:49:17.789813Z\", \"audit_version\": \"v3-dual-store+recovery-map\", \"measured_prose_w",
   "now": "{\"class\": \"full\", \"lacuna\": false, \"recovery_status\": \"COMPLETE\", \"residual_chars\": 22300, \"audited_at\": \"2026-07-17T04:49:17.789813Z\", \"audit_version\": \"v3-dual-store+recovery-map\", \"measured_prose_w"
  },
  {
   "date": "2026-08-05",
   "field": "description",
   "reason": "DW-??? intake (LABOR-prepared, TACHYON-verified: AXN match + factual probes vs record body)",
   "was": "The model collapse literature establishes that training generative models on their own outputs produces progressive distribution narrowing. The industry response is to seek human-written data as a corrective.",
   "now": "This paper challenges the binary classification of training data as either human-written and clean or machine-generated and contaminated. Its central distinction is that human authorship describes who typed the text, not whether the text’s distribution is independent of prior model mediation.\n\nThree mechanisms are proposed. First, sustained model use may reshape a writer’s unaided lexical, syntactic, and conceptual habits. Second, a multi-turn conversation creates a local feedback loop in which model vocabulary and framing enter later user turns. Third, users may learn accommodations specific to a particular model, causing inputs sent back to that model to carry its own interaction signature.\n\nThe paper replaces a binary contamination label with four continuous variables: habituation depth, turn position, task type, and model diversity. It predicts that later turns, creative tasks, heavy repeat users, and single-model use will carry stronger mediation signatures. It recommends contamination audits, turn-position weighting, multi-model diversification, tail-preservation benchmarks, and disclosure of chat-data use in model cards.\n\nThe record is unusually clear that the decisive studies have not been conducted. Its lexical-overlap and syntax estimates are illustrative. Four falsifiers specify large matched-user studies, turn-level corpus analysis, model-specific accommodation tests, and recursive training experiments using chat data.\n\nThe broad claim that providers increasingly treat user inputs as an owned clean-data source requires provider-specific policy and technical evidence. The paper does not establish that any named company trains on particular conversations or violates privacy. “The pristine source does not exist” is the paper’s theoretical conclusion, not a completed empirical audit."
  },
  {
   "date": "2026-09-05",
   "field": "related_deposits, remediation_note",
   "note": "Erratum linkage to #1577 (severity: evidence-status) per operator ruling of 2026-09-05; registry metadata only, canonical bytes and AXN untouched."
  }
 ],
 "date_modified": "2026-08-05",
 "publisher": "Pergamon Press",
 "journal_assignment": {
  "assigned": "2026-08-15",
  "by": "TACHYON under operator adjudication",
  "pass": 6,
  "method": "read per deposit — title and content_type, one at a time. No script classified anything.",
  "previous": "Machine-Mediated Reception Studies (MMRS)",
  "supersedes": "the 2026-06-21 preliminary batch mapping (#866), which assigned 864 deposits and put 371 in one venue",
  "authority": "data/cha-journals.json · datasets/venues/records/"
 },
 "related_deposits": [
  {
   "deposit_number": 1577,
   "axn": "AXN:0669.GOVERNANCE.🚀📜💙🪧↗️🔒",
   "relation": "corrected_by",
   "severity": "evidence-status",
   "note": "formal erratum — Ground 1 cites the Reverse Turing Test (#161) as establishing a result that #161 specifies but does not run; read as a proposed protocol, not a finding"
  }
 ],
 "remediation_note": "2026-09-05: ERRATUM #1577 (AXN:0669) — Ground 1's 'establishes' is an evidence-status error: #161 does not run the experiment, it specifies it. Registry-level cross-reference only; canonical bytes unchanged.",
 "falsification_conditions": "This paper's claims are falsifiable at the following points:\n\nF1. If a well-powered study (N > 500, pre-registered) demonstrates that heavy AI-chat users show no statistically significant tail-thinning in their unaided writing relative to matched non-users, Ground 1 fails.\n\nF2. If turn-level analysis of a large chat corpus (> 100,000 conversations) shows no within-conversation convergence of user inputs toward model outputs on lexical, syntactic, or perplexity measures, Ground 2 fails.\n\nF3. If users who interact exclusively with one model for six months produce inputs that are indistinguishable from inputs by users of other models (controlling for topic, length, and task type), Ground 3 fails.\n\nF4. If a training run using 50% chat data and 50% curated pre-2018 web data (predating GPT-2 and large-scale LLM deployment) produces no measurable model collapse over five generations of recursive training, the aggregate claim fails.\n\nNone of these studies has been conducted. This paper predicts their outcomes. The predictions are on the record.",
 "falsification_source": {
  "from": "body section",
  "heading": "## VI. Falsification Conditions",
  "harvested": "2026-09-08"
 },
 "line": "assembly-chorus",
 "line_parent": "method",
 "line_basis": "derived",
 "_projection": {
  "note": "Derived file. Canonical machine record is this entry in data/registry.json; the human record is the record_url. Do not edit this file.",
  "record_url": "https://www.alexanarch.org/s/records/856/",
  "self_url": "https://www.alexanarch.org/data/records/856.json",
  "registry_url": "https://www.alexanarch.org/data/registry.json",
  "text_url": "https://www.alexanarch.org/data/texts/AXN-0365-text.md",
  "oai_pmh": "https://www.alexanarch.org/oai?verb=Identify"
 }
}
