Certifying Regime Detectors Before Use
Value-at-risk forecasts are backtested before they are trusted; regime detectors, whose labels and triggers enter inference and allocation just as directly, face no analogous pre-use validation. We propose one. A certification protocol treats a frozen detector as a measurement instrument and, for a single pre-declared claim, measures five operating characteristics: power against a declared signal family (by controlled signal injection), false-alarm behavior under dependence-preserving resampled nulls, equivalence of outputs across calendar, volume, and event-time bars within a preregistered margin, detector-output information about the declared target, and the net value of that information under a declared cost convention. The protocol returns a certificate with a disposition and a gate-level failure profile, and refuses comparisons between detectors whose targets differ. We first certify a Holm-corrected variance-ratio cascade on BTCUSDT one-minute perpetual futures (2021–2025; n = 2,415,148 bars) and withhold the verdict, for measured reasons: the cascade cannot detect variance-ratio departures ≤ 0.10 anywhere on the preregistered 96-cell grid; its nominal 5% rule produced zero false alarms in 100 null replicates, a size distortion traceable to a mis-centered asymptotic reference; cross-clock equivalence certifies at q=2 but fails at the primary horizon q=5; and trigger information has a net value whose sign depends on the cost-attribution convention. We then show the diagnosis is actionable: recentering the decision statistic on the empirical null lowers the certified minimum detectable effect roughly tenfold — from 0.15 to 0.02 at q=2 — restores size control, and yields the protocol's first admissible certificate, for a bounded and now adequately powered exclusion of positive short-horizon serial dependence. Casting each gate as an e-value makes admission anytime-valid: one product certificate bounds the probability of ever falsely admitting a detector, under optional stopping, with no error budget across gates or monitoring times. A three-family exhibit (rolling-quantile, hidden-Markov, variance-ratio) yields cost-dominated, instrument-failed, and target-mismatched dispositions — failure reasons no performance comparison reports. Every value traces to frozen, version-pinned artifacts, with confirmatory and exploratory provenance separated per table cell.
- A pre-use certification protocol measuring five operating characteristics of a frozen detector against a single pre-declared claim
- An auditable certificate carrying a disposition and a gate-level failure profile, refusing comparisons between detectors whose targets differ
- An anytime-valid e-value formulation bounding the probability of ever falsely admitting a detector under optional stopping
- A fully reproducible artifact chain with confirmatory and exploratory provenance separated per table cell