EA-SEI-BCA-01 v2.0 is the submitted state of the Baseline Capture Architecture paper, deposited so that the record named in a journal's data-availability statement is the record a reader arrives at. It supersedes deposit #1453, which carried the architecture, the estimation theory and the accelerator feasibility argument but no worked demonstration.
The architecture's proposition is that a learned trigger, like any instrument, requires a control group. Where selection happens before storage and cannot be undone, the selection function can only be measured afterwards if something selection-independent was preserved at the time. The paper separates four channels usually conflated โ a content-independent probability sample taken before the audited selection, a full-rate tap of the representation the selector consumed, probability-weighted enrichment whose inclusion law is recorded, and shadow selectors whose decisions are logged without controlling acquisition โ and supplies the estimators that make retention of retrospectively defined classes a measurable quantity.
What v2.0 adds is the demonstration. On a synthetic stream of four million events with four injected anomaly classes, one of them constructed to lie exactly on the selector's learned manifold, four acquisition regimes are compared at matched storage budgets. The trigger-only archive audits itself as perfect, returning a retention of 1.000 for every class because the population it would need as a denominator no longer exists. A content-sensitive pseudo-baseline is confidently wrong rather than merely imprecise, estimating a shifted class at 0.76 with a tight interval against a true 0.016, because a second learned selector misses what the first one misses. The content-independent channel measures every class without selection bias, including a retention of exactly zero for the blind class, with an interval โ an estimate no content-sensitive channel in the experiment can produce at any budget. The rare class exhibits the paper's own sample-size relation, so the baseline states the limits of its own competence where the trigger archive cannot.
Two things are reported rather than tidied away: a miscalibrated enrichment design that starved the miss region, and a correction to the estimator lineage distinguishing Horvitz-Thompson unbiasedness for totals from the Hajek ratio estimator actually used for retention fractions.