docs: companion methods scaffold (BLIND) — RQ2-P3 mechanism + RQ3 churn
Draft the companion METHODS data-independent and blind, mirroring the lead-paper SS4 discipline: methods drafted, all Results/Discussion HELD-BLIND. No confirmatory run, no confirmatory number. - Two standalone methods sections (paper structure left OPEN — operator editorial call): (A) RQ2-P3 funnelling-mechanism study, header marked FREEZE-PENDING since its prereg §10 is unsigned — pool instrument, B×alpha manipulated-IV design, two-sided H1/H2/H3, run-level cluster bootstrap, re-worded §7 gate (cites stage-05-rq2p3-gate-clarification.md); (B) RQ3 companion from the frozen prereg §3/§4 — qwen2.5:3b + reproducibility caveat, churn kp=30/steps=20, frozen RQ3-P1-perf/latency/RQ3-P2 gates, Holm-7-supersedes note. - Intro/Related-Work deltas (unique-bridge/mix mechanism; churn-resilience + rebuild-fingerprint motivation), reusing only grounded lead citations plus Barton2025 from the frozen prereg's RQ3 DV. - The only numbers written are already-produced CALIBRATION gate values, each labelled calibration not result (RQ2-P3 dry preview mix; RQ3 churn-bites + classifier calibration). Results/Discussion are HELD-BLIND placeholders. HARD HOLDS: RQ2-P3 awaits §10 human freeze; RQ3 awaits operator-GO + live grid (added-latency is live-only). Frozen lead prereg f22331a72e… untouched; no HARKing; containment synthetic/offline, $0; worktree-only. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
This commit is contained in:
@@ -91,3 +91,6 @@ RQ2 RATIFIED (operator Andre, 2026-07-20, while blind) — observation-consisten
|
|||||||
- CHURN-BITES + CLASSIFIER GATES GREEN with real teeth (HARD HOLD satisfied): item1 churn bites — 1589 drops/1576 rebuilds, every cell fraction-with-drops=1.0; item2 rebuild-classifier AUC churned(kp30)-vs-baseline(kp5)=0.926≥0.90 SEPARABLE, baseline-vs-baseline null=0.518 blind; item3 agent cache-replay reproducible; item4 entropy=log2N. all_pass=True.
|
- CHURN-BITES + CLASSIFIER GATES GREEN with real teeth (HARD HOLD satisfied): item1 churn bites — 1589 drops/1576 rebuilds, every cell fraction-with-drops=1.0; item2 rebuild-classifier AUC churned(kp30)-vs-baseline(kp5)=0.926≥0.90 SEPARABLE, baseline-vs-baseline null=0.518 blind; item3 agent cache-replay reproducible; item4 entropy=log2N. all_pass=True.
|
||||||
- HARNESS FIX (not detector-tuning, logged): initial pooled-RAW-gap classifier landed AUC 0.779 (<0.90) — surfaced as candidate STOP, then diagnosed as a GROUPING-UNIT bug (pooling integer gaps floods AUC with ties). Switched calibration to PER-RUN mean-gap (the confirmatory grouping unit, cf. RQ2-P3) → 0.926 with a clean 0.518 null; frozen instrument (rebuild_interval_gaps, rebuild_classifier_auc) UNTOUCHED; rebuild-count rejected (teeth 1.0 but broken null 0.64). No fit to confirmatory data.
|
- HARNESS FIX (not detector-tuning, logged): initial pooled-RAW-gap classifier landed AUC 0.779 (<0.90) — surfaced as candidate STOP, then diagnosed as a GROUPING-UNIT bug (pooling integer gaps floods AUC with ties). Switched calibration to PER-RUN mean-gap (the confirmatory grouping unit, cf. RQ2-P3) → 0.926 with a clean 0.518 null; frozen instrument (rebuild_interval_gaps, rebuild_classifier_auc) UNTOUCHED; rebuild-count rejected (teeth 1.0 but broken null 0.64). No fit to confirmatory data.
|
||||||
- HARD HOLD REMAINS: NO RQ3 confirmatory battery run (run_rq3_battery live=True is operator+grid-gated; added-latency needs isolated-docker). Tests: tests/test_sor_rq3_wiring.py 9 passed; full SOR suite 194 passed (no regression). Lead prereg SHA f22331a72e… INTACT; RQ2-P3 §10 still empty; containment intact (synthetic/offline, $0 local); worktree-only.
|
- HARD HOLD REMAINS: NO RQ3 confirmatory battery run (run_rq3_battery live=True is operator+grid-gated; added-latency needs isolated-docker). Tests: tests/test_sor_rq3_wiring.py 9 passed; full SOR suite 194 passed (no regression). Lead prereg SHA f22331a72e… INTACT; RQ2-P3 §10 still empty; containment intact (synthetic/offline, $0 local); worktree-only.
|
||||||
|
- 2026-07-21 COMPANION METHODS drafted BLIND + committed (docs/stage-07-companion-methods.md), mirroring lead-paper SS4 discipline: two standalone methods sections — (A) RQ2-P3 mechanism (marked FREEZE-PENDING, §10 unsigned) + (B) RQ3 companion (frozen prereg §3/§4) — plus Intro/Related-Work deltas (unique-bridge/mix motivation; churn + rebuild-fingerprint motivation), reusing only grounded lead citations (+Barton2025 from the frozen prereg RQ3 DV).
|
||||||
|
- ALL Results/Discussion HELD-BLIND; the ONLY numbers written are already-produced CALIBRATION values, each labelled calibration NOT result (RQ2-P3 dry preview ρ 0→+0.838 / B=1 high-H; RQ3 gates churn-bites + classifier AUC≈0.93 sep / ≈0.52 null). Both confirmatory tracks stay behind their human gates: RQ2-P3 awaits §10 freeze SHA + sign-off; RQ3 awaits operator-GO + live isolated grid (added-latency is live-only).
|
||||||
|
- PAPER STRUCTURE LEFT OPEN (operator editorial call — second paper vs folded sections vs mechanism note; not pre-committed). No confirmatory run, no fabricated number, no HARKing; frozen lead prereg f22331a72e… UNTOUCHED; stray docs/.stage-07-paper-draft.md.swp deleted (not committed); containment synthetic/offline, $0; worktree-only feat/sor-consent-relay.
|
||||||
|
|||||||
@@ -0,0 +1,272 @@
|
|||||||
|
# Companion Methods (BLIND scaffold): The Unique-Bridge / Mix Mechanism (RQ2-P3) and Churn-Resilient Agent Selection (RQ3)
|
||||||
|
|
||||||
|
**Draft — companion methods, written BLIND. Results/Discussion HELD until each track clears its human gate.**
|
||||||
|
|
||||||
|
> **Paper-structure note (deliberately left OPEN).** Whether this material ships as a second
|
||||||
|
> standalone paper, as extension sections folded into the lead paper
|
||||||
|
> (`docs/stage-07-paper-draft.md`), or as a short mechanism note is an **operator editorial
|
||||||
|
> decision** and is **not** pre-committed here. The two methods tracks below are therefore
|
||||||
|
> written as **self-contained sections** that can be lifted into either structure.
|
||||||
|
>
|
||||||
|
> **Blinding & gating status (binding).**
|
||||||
|
> - **RQ2-P3 mechanism study** — the prereg (`docs/rq2p3-mechanism-prereg.md`, own slug
|
||||||
|
> `sor-consent-rq2p3`) is **NOT YET FROZEN** (§10 empty). Its methods section below is marked
|
||||||
|
> **FREEZE-PENDING**; **no confirmatory cell** runs until the operator records the SHA-256 in
|
||||||
|
> §10 and signs off.
|
||||||
|
> - **RQ3 companion** — hypotheses, gates, and analysis are **already frozen** in the lead prereg
|
||||||
|
> (`sor-consent-prereg.md`, SHA-256
|
||||||
|
> `f22331a72e0d0ccf38b787e63acabbe9d666456ec76076787a6d545c3193425b`, §3/§4/§6); the two open
|
||||||
|
> `[APPROVAL]` execution params were pinned blind in `docs/rq3-companion-run-brief.md`. **No
|
||||||
|
> confirmatory battery** runs until operator-GO on a live isolated grid (the added-latency DV is
|
||||||
|
> live-only).
|
||||||
|
> - **All Results / Discussion below are HELD-BLIND placeholders.** The **only** numbers written
|
||||||
|
> here are already-produced **calibration-gate** values, and every one is labelled *calibration*,
|
||||||
|
> never a confirmatory result. The frozen lead prereg is authoritative and **unedited**.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Abstract *(skeleton — HELD-BLIND; quantitative confirmatory claims withheld until each gate clears)*
|
||||||
|
|
||||||
|
The lead study measured a consent-gated, federated, nested-SSH relay instrument and reported two
|
||||||
|
honest non-confirmations: no measurable entry↔exit linkability leak (RQ1) and a Holm-significant
|
||||||
|
**shrink** of the per-circuit anonymity set under federation (RQ2-P1). This companion pursues the
|
||||||
|
two questions the lead paper could not close. **First (RQ2-P3′, a mechanism study):** the lead
|
||||||
|
"shrink" may be an **instrument artifact** — the bridge-federated topology assigns a *fresh*
|
||||||
|
willing bridge per circuit seed, so every adversary-observable exit signature is unique, every
|
||||||
|
anonymity set collapses to size 1, and entropy is driven to ≈0 by construction rather than by
|
||||||
|
funnelling. We re-instrument the willing-bridge layer as a **finite shared pool** with skewed
|
||||||
|
willingness and treat bridge **concentration as a manipulated independent variable** (a 3×3
|
||||||
|
dose-response over pool size and skew), asking **two-sided** whether concentration *reduces*
|
||||||
|
(funnel) or *raises* (mix) the anonymity set. **Second (RQ3, churn resilience):** we measure
|
||||||
|
whether a **local open-weight agent** path-selector (`qwen2.5:3b`) retains throughput and adds
|
||||||
|
tolerable latency under a pinned churn schedule, without leaving a classifiable **rebuild
|
||||||
|
fingerprint**. Both tracks are pre-registered, detector-frozen, and calibration-gated before any
|
||||||
|
confirmatory cell. *(Confirmatory findings withheld: RQ2-P3 pending freeze; RQ3 pending
|
||||||
|
operator-GO on the live grid.)*
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 1. Introduction (deltas beyond the lead paper)
|
||||||
|
|
||||||
|
The lead paper (G4 + RQ1 + RQ2) established the consent-gate instrument and reported its
|
||||||
|
linkability and anonymity-set readings. Two threads there were *raised but not resolved*, and this
|
||||||
|
companion is scoped to exactly those.
|
||||||
|
|
||||||
|
**(a) The unique-bridge / mix mechanism.** The lead RQ2-P1 result — federation **shrinks** the
|
||||||
|
anonymity set (ΔH < 0) — was reported honestly, but the lead paper also flagged its RQ2-P3
|
||||||
|
mechanism test as **degenerate as-instrumented**: the bridge-federated arm assigns a fresh bridge
|
||||||
|
per circuit seed, so top-3 bridge concentration is a constant `c_i = 1/C` with **zero variance**
|
||||||
|
and Spearman ρ is undefined. The mechanistic reading (developed in
|
||||||
|
`docs/note-unique-bridge-artifact.md`) is that the adversary's observable is an
|
||||||
|
`exit_signature = (exit_house, bridge_label)`; a unique bridge per circuit makes every signature
|
||||||
|
unique, so the observation-consistent anonymity set is size 1 and H≈0 **by injective construction,
|
||||||
|
not by funnelling**. If that is right, a *finite shared* bridge pool should make circuits share
|
||||||
|
signatures, enlarge the anonymity set, and act as a **mix** (concentration *raises* H) — the
|
||||||
|
**opposite** of the naive funnel intuition. This makes RQ2-P3′ a test of whether the lead
|
||||||
|
"shrink" headline is a unique-bridge artifact that a mechanism study can qualify or correct. This
|
||||||
|
mix reading connects the consent-gate bridge to the classical mix [Chaum1981] and to
|
||||||
|
information-theoretic set metrics [Serjantov2002; Diaz2002] the lead paper already adopts.
|
||||||
|
|
||||||
|
**(b) Churn-resilient agent selection.** The lead paper held the selector at `static`; RQ3 asks
|
||||||
|
whether an **adaptive** selector improves resilience when the relay pool churns. Two costs bound
|
||||||
|
any such gain and are the confirmatory tension: (i) rebuilding a circuit after a dropped hop adds
|
||||||
|
latency and can erode throughput, and (ii) the *timing pattern* of rebuilds is itself a
|
||||||
|
side-channel — a rebuild-event classifier could fingerprint the selector, echoing website- and
|
||||||
|
flow-fingerprinting results on onion transports [SirinamIJW18; RahmanSMGW20] and the
|
||||||
|
rebuild/timing-classifier spirit of CLASI [Barton2025], and compounding statistical-disclosure
|
||||||
|
exposure over repeated circuits [Danezis2003]. RQ3 therefore pairs a **performance** gate with an
|
||||||
|
**anonymity** (non-fingerprint) gate: an agent selector only "helps" if it retains throughput at
|
||||||
|
tolerable added latency **and** its rebuild pattern is not classifiable.
|
||||||
|
|
||||||
|
## 2. Related work (deltas)
|
||||||
|
|
||||||
|
The lead paper's Related Work (onion routing / SOR [Egners2012], flow correlation
|
||||||
|
[NasrBH18; OhYMH22], anonymity metrics [Serjantov2002; Diaz2002], social-trust G4 neighbours) is
|
||||||
|
inherited unchanged. The companion adds two narrow deltas, citing **only** already-grounded
|
||||||
|
references:
|
||||||
|
|
||||||
|
- **Bridge-as-mix vs. bridge-as-funnel.** Concentrating flows through few willing bridges can be
|
||||||
|
read either as a funnel (fewer distinct observation classes → smaller sets) or as a mix
|
||||||
|
[Chaum1981] (shared observation class → larger sets). The set-size effect is quantified with the
|
||||||
|
same entropy metrics the lead paper uses [Serjantov2002; Diaz2002]; the companion's contribution
|
||||||
|
is a **manipulated-concentration dose-response** that adjudicates the sign, not a new estimator.
|
||||||
|
- **Rebuild-timing as a fingerprint.** Churn-driven circuit rebuilds create a timing series an
|
||||||
|
adversary may classify; this is the fingerprinting/timing lineage [SirinamIJW18; RahmanSMGW20;
|
||||||
|
Barton2025] applied to *selector-induced* rebuild events rather than page loads. The companion
|
||||||
|
adopts a **frozen, fixture-calibrated** rebuild classifier and reads its AUC as an instrument
|
||||||
|
reading, mirroring the lead paper's frozen-correlator discipline (no correlator/classifier
|
||||||
|
state-of-the-art is claimed).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 3. Methods A — RQ2-P3′ funnelling-mechanism study **[FREEZE-PENDING (prereg §10 not yet signed)]**
|
||||||
|
|
||||||
|
> **This section describes a study whose prereg (`docs/rq2p3-mechanism-prereg.md`) is NOT frozen.**
|
||||||
|
> Design and parameters are operator-approved and locked; the **human freeze checkpoint** (SHA-256
|
||||||
|
> in §10 + sign-off) is outstanding. No confirmatory cell is run until then. Everything below is
|
||||||
|
> the pre-registered *plan*; the calibration numbers cited are the pre-registered §7 gate output,
|
||||||
|
> explicitly **not** confirmatory results.
|
||||||
|
|
||||||
|
**Design (manipulated-IV dose-response).** A new assembler topology, `bridge-federated-pool`,
|
||||||
|
replaces the fresh-per-seed bridge with a **finite willing-bridge pool** of size `B` under a fixed
|
||||||
|
Zipf willingness skew `alpha`: `weights = zipf_weights(B, alpha)` derived from the **cell** seed
|
||||||
|
(so the willingness profile is fixed within a run), and each circuit draws
|
||||||
|
`idx = weighted_draw(sha256("sor-bridge-pool|{circuit_seed}"), weights)` → `bridge#{idx:02d}`,
|
||||||
|
a label **reused** across circuits so concentration genuinely varies. Everything downstream (hop
|
||||||
|
structure, houses, exit-signature grouping, Miller–Madow entropy, BCa bootstrap) is **identical**
|
||||||
|
to the frozen lead pipeline; the lead `bridge-federated` branch is **untouched and bit-reproducible**.
|
||||||
|
The manipulation grid is **B ∈ {2, 4, 8} × alpha ∈ {0 (uniform), 1.0, 2.0}** = 9 concentration
|
||||||
|
cells; run order randomized within cell from an ordering seed distinct from the data seeds.
|
||||||
|
|
||||||
|
**Hypotheses (two-sided; direction not presumed).**
|
||||||
|
- **H1 (within-cell association).** Spearman ρ between per-circuit top-3 willing-bridge
|
||||||
|
concentration `c_i` and per-circuit entropy `H_i`. **Funnel** iff BCa 95% CI < 0; **mix** iff CI
|
||||||
|
> 0; **inconclusive** iff CI spans 0.
|
||||||
|
- **H2 (dose-response).** OLS slope β of per-run mean-H on per-run mean top-3 concentration over
|
||||||
|
9 cells × 30 runs = 270 clustered points; cell-level BCa CI; funnel iff slope CI < 0, mix iff CI > 0.
|
||||||
|
- **H3 (joint, direction-agnostic).** Mechanism **RESOLVED** iff H1 and H2 agree in sign and both
|
||||||
|
exclude 0 — the *sign* (funnel vs mix) is the finding; **unresolved** if either spans 0.
|
||||||
|
|
||||||
|
**Dependent variables.** Per-circuit `H_i` (Miller–Madow entropy of the uniform posterior over the
|
||||||
|
observation-consistent anonymity set, inherited verbatim) and per-circuit top-3 concentration `c_i`
|
||||||
|
(`confirm_load_rq2.bridge_concentration`, unchanged).
|
||||||
|
|
||||||
|
**Sampling.** R = 30 seeded runs/cell, C = 50 circuits/run (matched to the lead study); base seed
|
||||||
|
S0 = 20260719, per-cell seed `SHA256(S0 ‖ cell_id ‖ run_index)`; fixed stopping rule (all 9 cells ×
|
||||||
|
R to completion; uninformative cell → inconclusive; no optional stopping).
|
||||||
|
|
||||||
|
**Analysis.** Effect size + BCa 95% CI (10,000 resamples, α = 0.05) for every test; **run-level
|
||||||
|
cluster bootstrap** (resample whole runs, not circuits) because circuits sharing a bridge have
|
||||||
|
identical `c_i` and correlated `H_i` — the same pseudo-replication defect the lead paper flagged
|
||||||
|
for RQ1-P1. Holm–Bonferroni over **this study's own family** {H1-pooled, H2-slope}; the lead
|
||||||
|
family-of-7 is closed and **not** reopened here. Any per-cell ρ contrast or ΔH replication is
|
||||||
|
labelled **EXPLORATORY**, never a re-run of the frozen RQ2-P1.
|
||||||
|
|
||||||
|
**Instrument-validation gate (§7, re-worded pre-freeze — cite
|
||||||
|
`docs/stage-05-rq2p3-gate-clarification.md`).** The §7 items were **re-worded before freeze**
|
||||||
|
because the original items 1–2 encoded the naive-funnel prior and were mechanically wrong under the
|
||||||
|
ratified posterior (transparent deviation logged; no hypothesis changed — H1/H2/H3 stay two-sided).
|
||||||
|
The re-worded gate validates the **instrument**, not a sign:
|
||||||
|
1. the **frozen** `bridge-federated` branch (not the pool) still shows the lead degeneracy — unique
|
||||||
|
signatures → `m_i = 1` → `H_i ≈ 0`, constant `c_i = 1/C` (a pool draws with replacement and
|
||||||
|
*cannot* reproduce the injective fresh-bridge degeneracy, so the regression teeth live on the
|
||||||
|
untouched branch);
|
||||||
|
2. the **B = 1 boundary** yields `c = 1.0` (concentration tooth) **and**, under the ratified
|
||||||
|
posterior, `H` at the **high** end (maximal mix) — the naive "low H" gloss is refuted by
|
||||||
|
construction;
|
||||||
|
3. realized mean top-3 concentration is **monotone** (decreasing in B, increasing in alpha);
|
||||||
|
4. entropy calibration inherited (H = log₂N on equiprobable synthetic senders).
|
||||||
|
A **§7 scope note** records that the gate **must not** pre-assert the H-vs-concentration sign —
|
||||||
|
that sign *is* the two-sided confirmatory question; baking it in would be funnel-circular.
|
||||||
|
|
||||||
|
**Pre-registered calibration finding (NOT a confirmatory result).** The dry §7 pass — synthetic,
|
||||||
|
offline, no confirmatory record read — already **previews a mix**: across the sweep Spearman ρ runs
|
||||||
|
from ≈0 up to **+0.838** (all cells ρ ≥ 0), the B = 1 boundary sits at high entropy (≈2.54 bits vs
|
||||||
|
the fresh-bridge reference ≈0.0), and monotonicity + entropy calibration pass. This is surfaced
|
||||||
|
**openly as a pre-registered calibration preview**, per the §7 scope note; it does **not** relax the
|
||||||
|
two-sided pre-commitment, and the confirmatory sign remains withheld until after freeze. Honest
|
||||||
|
disclosure the eventual write-up must carry: because the dry pass already previews the mix
|
||||||
|
direction, the confirmatory battery **quantifies a dose-response already visible at calibration**;
|
||||||
|
the two-sided pre-commitment stands and the lead RQ2-P1 headline is not re-litigated.
|
||||||
|
|
||||||
|
## 4. Methods B — RQ3 churn-resilient agent selector (frozen prereg)
|
||||||
|
|
||||||
|
> Hypotheses, gates, DVs, and analysis are **frozen** in `sor-consent-prereg.md` (§3/§4/§6) and are
|
||||||
|
> restated, not redefined. The two open `[APPROVAL]` execution params were pinned **blind** to RQ3
|
||||||
|
> outcomes (`docs/rq3-companion-run-brief.md` §2).
|
||||||
|
|
||||||
|
**Design.** Selector strategy {`static`, `random`, `agent`} at the RQ3 control cell
|
||||||
|
(single-house / bridge-off) under a pinned churn schedule; `static` is the interleaved control, and
|
||||||
|
control runs are bracketed before and after the {`random`, `agent`} treatments to catch grid drift.
|
||||||
|
The cells are enumerated **separately** from the frozen 6-cell lead lattice so the lead battery
|
||||||
|
stays bit-reproducible.
|
||||||
|
|
||||||
|
**Pinned execution parameters (blind).**
|
||||||
|
- **Agent = `qwen2.5:3b`** via local Ollama (`agent_selector.OllamaAgentPolicy`, temperature 0,
|
||||||
|
per-run seed, `(seed, state-hash)` decision cache, deterministic heuristic fallback on query
|
||||||
|
failure). Local / open-weight, **$0**; the Claude/frontier arm (`ClaudeExploratoryPolicy`) stays
|
||||||
|
**inert / EXPLORATORY / budget-gated** and is not wired.
|
||||||
|
- **Reproducibility caveat (accepted; must be stated in the paper).** Ollama at temperature 0 is
|
||||||
|
**not bit-identical across machines** (quantization / GPU logit drift). The agent arm is
|
||||||
|
reproducible via the **committed decision-log + `(seed, state-hash)` cache replay**, *not* via
|
||||||
|
independent model re-execution on other hardware — the same honesty class as the RQ1 timing
|
||||||
|
caveat. The decision log + cache are committed as the reproducibility anchor.
|
||||||
|
- **Churn = `kill_prob_pct = 30`, `steps = 20`**, one deterministic schedule per run seeded from
|
||||||
|
the same `SHA256(S0 ‖ cell ‖ run)` family; low-churn calibration baseline `kill_prob_pct = 5`.
|
||||||
|
|
||||||
|
**Dependent variables (frozen).** Throughput retention (throughput under churn / no-churn
|
||||||
|
baseline); added latency = median end-to-end latency(agent) − median latency(best baseline arm), in
|
||||||
|
ms — a **live** measurement only; and rebuild-classifier AUC over the rebuild-event time series (the
|
||||||
|
per-run mean inter-rebuild-gap signal), per the [Barton2025] CLASI spirit.
|
||||||
|
|
||||||
|
**Confirmatory gates (frozen, family-of-7).**
|
||||||
|
|
||||||
|
| Test | Frozen gate |
|
||||||
|
|---|---|
|
||||||
|
| RQ3-P1-perf | throughput-retention(agent) − max(static, random): 95% CI lower bound **≥ 10 pp** |
|
||||||
|
| RQ3-P1-latency | added-latency(agent): 95% CI upper bound **≤ 100 ms** |
|
||||||
|
| RQ3-P2 | rebuild-classifier AUC: 95% CI upper bound **≤ 0.60** |
|
||||||
|
| RQ3-P3 | logical AND: **CONFIRM** iff P1 ∧ P2 (perf gain *without* a rebuild fingerprint); else H0 |
|
||||||
|
|
||||||
|
R = 30 runs/cell, C = 50 circuits/run, fixed stopping rule (inherited unchanged).
|
||||||
|
|
||||||
|
**Analysis + multiplicity (Holm-7 supersedes note).** The three RQ3 tests were always in the
|
||||||
|
frozen size-7 family {RQ1-P1, RQ1-P2, RQ2-P1, RQ2-P3, RQ3-P1-perf, RQ3-P1-latency, RQ3-P2}. Once
|
||||||
|
all seven p-values exist, the companion computes the **exact Holm-7** step-down over the whole
|
||||||
|
family; this is the **authoritative** final correction and **supersedes** the lead paper's
|
||||||
|
deliberately conservative *partial* embedding (7/6/5/4 report-4) — both remain valid, the partial
|
||||||
|
never under-corrects, and RQ1-P1 / RQ2-P1 survive regardless (raw p ≈ 0). Effect size + BCa 95% CI
|
||||||
|
for every test; nulls reported honestly (a selector that does **not** beat baselines, or a rebuild
|
||||||
|
pattern that **is** classifiable, is the finding). The QUIC / `ssh3` transport arm stays
|
||||||
|
**EXPLORATORY and deferred** (design decision D3), never in the Holm family.
|
||||||
|
|
||||||
|
**Calibration gates (already green; NOT confirmatory results).** Two boolean gates block the RQ3
|
||||||
|
confirmatory battery and have both passed on a **dry, synthetic, offline** pass:
|
||||||
|
- **Churn-bites** — at the pinned `kp = 30 / steps = 20` the churn genuinely bites (non-zero
|
||||||
|
drops/rebuilds across every RQ3 cell), so the retention and classifier tests are not degenerate.
|
||||||
|
- **Rebuild-classifier calibration** — churned (`kp = 30`) vs low-churn baseline (`kp = 5`) is
|
||||||
|
**separable** on the per-run mean inter-rebuild-gap signal (calibration AUC ≈ **0.93**), while
|
||||||
|
baseline-vs-baseline is **not** (null AUC ≈ **0.52**); plus an agent cache-replay reproducibility
|
||||||
|
check and the inherited entropy calibration. These are **calibration** readings on labelled
|
||||||
|
control signals, **not fit to confirmatory cells**; the frozen instrument
|
||||||
|
(`rebuild_interval_gaps`, `rebuild_classifier_auc`) is unchanged.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 5. Results *(HELD-BLIND — no confirmatory number written)*
|
||||||
|
|
||||||
|
- **RQ2-P3′ (H1 / H2 / H3).** *HELD — pending prereg freeze (§10) and the confirmatory run.* No
|
||||||
|
Spearman ρ, OLS slope, CI, or Holm decision is computed or inspected until after freeze. *(The
|
||||||
|
only figures on record are the §4/§3-A pre-registered **calibration** preview — ρ 0→+0.838, B=1
|
||||||
|
high-H — explicitly labelled calibration, not a result.)*
|
||||||
|
- **RQ3-P1-perf / RQ3-P1-latency / RQ3-P2 / RQ3-P3.** *HELD — pending operator-GO on the live
|
||||||
|
isolated grid.* Throughput-retention margin, added-latency (live-only), and rebuild-classifier
|
||||||
|
AUC over confirmatory cells are **not** computed; the only figures on record are the two green
|
||||||
|
**calibration** gates (churn-bites; classifier AUC ≈0.93 sep / ≈0.52 null), labelled calibration.
|
||||||
|
- **Holm-7 (companion, authoritative).** *HELD* until all seven confirmatory p-values exist.
|
||||||
|
|
||||||
|
## 6. Discussion *(HELD-BLIND — placeholders)*
|
||||||
|
|
||||||
|
- **Does a shared bridge funnel or mix?** *HELD.* If H1/H2 confirm **mix** (ρ, slope > 0), the
|
||||||
|
companion qualifies the lead "shrink" as partly a unique-bridge artifact; if **funnel** (< 0), the
|
||||||
|
lead reading is mechanistically corroborated; if inconclusive, the mechanism stays open. The sign
|
||||||
|
is the finding; the lead RQ2-P1 result is reported as-is and not re-litigated.
|
||||||
|
- **Does the agent selector help without leaking?** *HELD.* RQ3-P3 confirms only if the agent both
|
||||||
|
clears the +10 pp retention / ≤100 ms latency bar **and** leaves a non-classifiable rebuild
|
||||||
|
pattern (AUC CI upper ≤ 0.60); either failure is an honest H0 (perf cancelled by teardown
|
||||||
|
overhead, or the rebuild timing is a usable fingerprint).
|
||||||
|
- **Scope & limitations.** *HELD* — will inherit the lead paper's lab-grid external-validity scope,
|
||||||
|
add the RQ2-P3 as-instrumented concentration/mix caveat, and carry the RQ3 agent-reproducibility
|
||||||
|
caveat (cache-replay, not cross-hardware re-execution) with equal prominence.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## References
|
||||||
|
|
||||||
|
Inherit the lead paper's reference list (`docs/stage-07-paper-draft.md` §References) unchanged. The
|
||||||
|
companion cites only references already grounded in the frozen sources: **Chaum1981, Serjantov2002,
|
||||||
|
Diaz2002, Egners2012, NasrBH18, OhYMH22, SirinamIJW18, RahmanSMGW20, Danezis2003** are in the lead
|
||||||
|
paper's list; **Barton2025** (CLASI rebuild/timing-classifier spirit) is grounded in the frozen
|
||||||
|
lead prereg's RQ3 dependent-variable definition (`sor-consent-prereg.md` §3) and carries into the
|
||||||
|
companion's assembled list. No reference is added that cannot be grounded from the frozen prereg or
|
||||||
|
stage-01 literature.
|
||||||
Reference in New Issue
Block a user