Authoritative identifier:AXN:0351.EMPIRICAL.🟢🏗️🙏📜🔔🌱 — content-derived; resolves at /s/axn/0351/. Cite this record by AXN and URL.
former DOI10.5281/zenodo.20722522 — severed 2026-06-19 by the registrant; resolves to a tombstone (HTTP 410); resolution page
former DOI10.5281/zenodo.20722523 — severed 2026-06-19 by the registrant; resolves to a tombstone (HTTP 410); resolution page
Description
This document is a living work plan for extracting and governing the Crimson Hexagonal Archive’s coined vocabulary. Its initial phase templates say “not started,” but the later progress table and session log record substantial execution. It should therefore be read as a plan whose headings preserve the original design while its ledger records the state reached during the same work cycle.
The planned workflow has five phases. Metadata are pulled from every repository record and candidate terms extracted from titles, descriptions, keywords, codes, quotations, and emphasis. File bodies are then processed in resumable batches. A human review removes ordinary-language false positives, restores missed coinages, assigns definitions and categories, and reconciles known institutional lists. The resulting JSON and Markdown index are deposited and surfaced as a searchable table. High-priority terms may later receive provenance packets.
The progress ledger reports 845 metadata records processed, 1,524 repeated candidate terms, tiered canonicalization, 735 of 800 downloadable body records processed, sixty-five download failures, and 129 of 131 capture-registry queries cross-referenced. Merge, noise filtering, human review, and the final v1.0 index remained pending, while an initial raw-data deposit was in progress.
This is a valuable continuity artifact because it makes unfinished work visible. It is not itself the final index, and the extracted counts are not counts of confirmed coinages. Automated phrase extraction will include names, ordinary phrases, bibliographic language, duplicates, and template artifacts. The internal instructions referring to /home/claude/, session compaction, and present_files are ephemeral execution notes and should not be treated as public architectural requirements.
Wiki Article
Crimson Hexagonal Archive: Term Index Work Plan is a living plan and progress ledger for building a canonical vocabulary index.
It describes metadata extraction, body-text processing, human review, cross-reference work, publication, versioning, and optional provenance-packet generation.
Although the original phase headings begin as “not started,” the later ledger records extensive completion of metadata and body processing. Merge, human validation, and the final index were still pending.
The candidate-term totals are automated extraction results rather than a verified count of archive coinages.
Also published as a standalone entry: /s/wiki/836/
Concepts Defined
Progress checkpoint[structural] After Phase 1, we have ~60-70% of coinages from metadata alone. Save all three files to /home/claude/ and present. If session compacts here, next session loads these files and proc
Session 1 (16 June 2026)[structural] Work plan created. Phase 1.1 complete (845 records pulled). Phase 1.2 complete (1,524 terms extracted with count>=2; 5,951 total unique keywords). Phase 1.3 (canonicalization and c
Full Text
CRIMSON HEXAGONAL ARCHIVE: TERM INDEX WORK PLAN
EA-REGISTRY-TERMINDEX-PLAN v1.0
Author: Lee Sharks (ORCID 0009-0000-1599-0703)
Date: 16 June 2026
Status: Work plan with progress tracking
Purpose: Systematic extraction, canonicalization, and versioning of all coined terms, concepts, entities, frameworks, operators, institutions, heteronyms, and designations across the Crimson Hexagonal Archive (~841+ deposits)
Rationale
The archive is minting terms faster than they imprint to long-term memory. An estimated 200-300 coined concepts exist across 841+ deposits without a unified index. Once built, the index becomes:
A retrieval surface (each term becomes a searchable node)
A density map (which terms cluster, which are orphaned)
A versioning base (new deposits append; the initial build is the hard part)
An MPAI generator (each indexed term can receive a provenance coupling kernel)
A measure of the archive's actual vocabulary (what it has coined vs. what it thinks it has coined)
Phase 1: Metadata Pull
Status: NOT STARTED
Estimated compute: 30-45 minutes
Resumable: Yes — paginate via Zenodo API, save after each page
1.1 Pull all records from crimsonhexagonal community
API endpoint: https://zenodo.org/api/records?communities=crimsonhexagonal&size=200&page=N
Expected: ~841 records across 5 pages
Save: JSON file with record ID, DOI, title, description, keywords, creators, publication_date, version, related_identifiers
Output: termindex-metadata-raw.json
1.2 Extract terms from metadata fields
Parse each record's title, keywords, and description
Progress checkpoint: After Phase 1, we have ~60-70% of coinages from metadata alone. Save all three files to /home/claude/ and present. If session compacts here, next session loads these files and proceeds to Phase 2.
Session 1 (16 June 2026): Work plan created. Phase 1.1 complete (845 records pulled). Phase 1.2 complete (1,524 terms extracted with count>=2; 5,951 total unique keywords). Phase 1.3 (canonicalization and categorization) ready for next session or human review. Key finding: the archive has 6,256 keyword instances across 845 records, with the top terms being Crimson Hexagonal Archive (439), semantic economy (267), Crimson hexagon (248), distributed epic (154), NH-OS (149), operative semiotics (124), training layer literature (121). The API paginates at max size=25, requiring 34 pages. The metadata-raw and metadata-terms JSON files are the continuity artifacts for the next session.
2026-08-01 — journal: Wave 6 venue normalization: full canonical journal name per MANUS ruling 2026-08-01 (venues.json authority)
2026-08-04 — publisher: PUB-POPULATE: dc:publisher from venues.json v1.1 press mapping (CP-R3 RULED-EXTENDED 2026-08-01); Alexanarch = publisher of record where no imprint applies
2026-08-04 — status: W12 STATUS-VOCABULARY v1.0 (MANUS ratified 2026-08-04): controlled vocabulary {ACTIVE, SUPERSEDED, WITHDRAWN, DRAFT}; MINTED_UNREVIEWED false on a 100%-audited corpus; freetext annotations preserved losslessly in body_status.status_note