paper: fill RQ3 half of combined companion (post-seal, both tracks un-blinded)

Replace the HELD-BLIND RQ3 Results/Discussion + authoritative-Holm-7
placeholders in docs/stage-07-companion-methods.md with real numbers read
from the sealed record only (rq3-confirmatory-analysis.json e09c66ef…):

  RQ3-P1-perf   -0.6pp  CI[-1.58,+0.39]pp   -> H0 (no +10pp agent margin)
  RQ3-P1-latency -13.5ms CI[-52.1,+34.9]ms  -> within 100ms budget
  RQ3-P2 AUC     0.587  CI[0.458,0.703]      -> fingerprint NOT excluded
  RQ3-P3 = H0 (perf fails AND P2 fails)

Authoritative Holm-7 (family size 7, supersedes lead partial embedding, D3):
survivors RQ1-P1 / RQ2-P1 / RQ2-P3; non-survivors RQ1-P2 / RQ3-P2 /
RQ3-P1-perf / RQ3-P1-latency. RQ2-P3 slot carries the mechanism-corrected
primary p, not the lead degenerate p=1.

Calibration disclosure carried: the 0.93 RQ3 calibration AUC is regime
discrimination (churn kp30 vs kp5), distinct from the 0.587 confirmatory
selector-vs-selector fingerprint — no HARKing, detectors frozen.

RQ2-P3 half not re-litigated (kept as filled at e0d865d). Nulls reported
openly as results. Frozen lead prereg f22331a72e… and RQ2-P3 prereg
8db4e8a7… both intact. $0/offline, worktree-only.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
This commit is contained in:
leetcrypt
2026-07-22 08:53:50 -07:00
parent 47faee3a92
commit 1b6449b827
2 changed files with 84 additions and 26 deletions
+1
View File
@@ -109,3 +109,4 @@ RQ2 RATIFIED (operator Andre, 2026-07-20, while blind) — observation-consisten
- 2026-07-22 RQ3 BATTERY COMPLETE (90/90 runs, live-docker-e2e, exited clean) → UN-BLIND + FROZEN analysis. RAW SEALED: `.../20260722T040640Z/confirmatory-data/SHA256SUMS` (battery-results.json `5b61e461…` + 90 rq3-run.json sidecars). New analyzer `analysis/rq3_confirm.py` (RQ3 analogue of rq2p3_confirm; frozen `stats`/`metrics` UNTOUCHED — run-level multi-arm bootstrap mirrors `two_sample_diff_ci`, 10k BCa, α=0.05). Sealed analysis `.../20260722T040640Z/analysis/rq3-confirmatory-analysis.json` (SHA `e09c66ef…`). Tests `tests/test_sor_rq3_confirm.py` 6 passed; full SOR suite 207 passed (no regression). - 2026-07-22 RQ3 BATTERY COMPLETE (90/90 runs, live-docker-e2e, exited clean) → UN-BLIND + FROZEN analysis. RAW SEALED: `.../20260722T040640Z/confirmatory-data/SHA256SUMS` (battery-results.json `5b61e461…` + 90 rq3-run.json sidecars). New analyzer `analysis/rq3_confirm.py` (RQ3 analogue of rq2p3_confirm; frozen `stats`/`metrics` UNTOUCHED — run-level multi-arm bootstrap mirrors `two_sample_diff_ci`, 10k BCa, α=0.05). Sealed analysis `.../20260722T040640Z/analysis/rq3-confirmatory-analysis.json` (SHA `e09c66ef…`). Tests `tests/test_sor_rq3_confirm.py` 6 passed; full SOR suite 207 passed (no regression).
- RQ3 RESULT (effect+CI, never bare p): **RQ3-P1-perf FAIL/H0** — retention margin agentmax(static,random) = **0.6pp**, BCa CI [1.58pp, +0.39pp]; every selector heals ~all churn drops (~99% retention) so no ≥+10pp agent gain. **RQ3-P1-latency HOLDS budget** — added-latency(agentmin-baseline=random) = **13.5ms**, CI [52.1, +34.9]ms, upper ≤100ms (agent not slower). **RQ3-P2 FAIL/not-excluded** — rebuild-classifier AUC(agent vs pooled baseline, per-run mean-gap) = **0.587**, CI [0.458, 0.703], upper 0.703 > 0.60 → fingerprint NOT excluded (underpowered at n=30). **RQ3-P3 = H0** (P1-perf fails ∧ P2 fails). Honest null: the local agent selector neither beats baselines nor is certifiably non-classifiable. - RQ3 RESULT (effect+CI, never bare p): **RQ3-P1-perf FAIL/H0** — retention margin agentmax(static,random) = **0.6pp**, BCa CI [1.58pp, +0.39pp]; every selector heals ~all churn drops (~99% retention) so no ≥+10pp agent gain. **RQ3-P1-latency HOLDS budget** — added-latency(agentmin-baseline=random) = **13.5ms**, CI [52.1, +34.9]ms, upper ≤100ms (agent not slower). **RQ3-P2 FAIL/not-excluded** — rebuild-classifier AUC(agent vs pooled baseline, per-run mean-gap) = **0.587**, CI [0.458, 0.703], upper 0.703 > 0.60 → fingerprint NOT excluded (underpowered at n=30). **RQ3-P3 = H0** (P1-perf fails ∧ P2 fails). Honest null: the local agent selector neither beats baselines nor is certifiably non-classifiable.
- AUTHORITATIVE HOLM-7 (frozen family size=7; supersedes lead's conservative partial embedding, D3). RQ2-P3 slot carries the mechanism-corrected primary **H1-pooled Spearman p=0** (not the lead's degenerate p=1). **SURVIVORS: RQ1-P1 (r1×7), RQ2-P1 shrink (r2×6), RQ2-P3 mix (r3×5)** all holm_p=0. NON-survivors: RQ1-P2 (r4×4, holm_p=0.365), RQ3-P2 (r5×3, 0.511), RQ3-P1-perf (r6×2, 0.511), RQ3-P1-latency (r7×1, 0.511). Lead RQ1-P1 & RQ2-P1 survive regardless. Both prereg SHAs intact; $0/offline; worktree-only. NEXT: fill RQ3 half of companion paper (stone 2). - AUTHORITATIVE HOLM-7 (frozen family size=7; supersedes lead's conservative partial embedding, D3). RQ2-P3 slot carries the mechanism-corrected primary **H1-pooled Spearman p=0** (not the lead's degenerate p=1). **SURVIVORS: RQ1-P1 (r1×7), RQ2-P1 shrink (r2×6), RQ2-P3 mix (r3×5)** all holm_p=0. NON-survivors: RQ1-P2 (r4×4, holm_p=0.365), RQ3-P2 (r5×3, 0.511), RQ3-P1-perf (r6×2, 0.511), RQ3-P1-latency (r7×1, 0.511). Lead RQ1-P1 & RQ2-P1 survive regardless. Both prereg SHAs intact; $0/offline; worktree-only. NEXT: fill RQ3 half of companion paper (stone 2).
- 2026-07-22 COMPANION PAPER — RQ3 HALF FILLED POST-SEAL (stone 2, D3 combined paper). `docs/stage-07-companion-methods.md`: §5 Results + §6 Discussion RQ3 placeholders AND the authoritative-Holm-7 placeholder REPLACED with real numbers FROM THE SEALED RECORD ONLY (perf 0.6pp CI[1.58,+0.39]pp H0; latency 13.5ms CI[52.1,+34.9] within budget, min-baseline=random; P2 AUC 0.587 CI[0.458,0.703] fingerprint-not-excluded; P3=H0). Authoritative Holm-7 in-paper: survivors RQ1-P1/RQ2-P1/RQ2-P3, non-survivors RQ1-P2/RQ3-P2/RQ3-P1-perf/RQ3-P1-latency. Abstract + top blinding block un-blinded (both tracks). RQ2-P3 half NOT re-litigated (kept as filled at e0d865d). DISCLOSURE carried: 0.93 RQ3 calibration AUC = regime discrimination (churn kp30 vs kp5), DISTINCT from the 0.587 confirmatory selector-vs-selector fingerprint — no HARKing, gates frozen. $0/offline, worktree-only; both prereg SHAs intact. NEXT: report + HOLD for review.
+83 -26
View File
@@ -1,6 +1,6 @@
# Companion Methods (BLIND scaffold): The Unique-Bridge / Mix Mechanism (RQ2-P3) and Churn-Resilient Agent Selection (RQ3) # Companion Methods (BLIND scaffold): The Unique-Bridge / Mix Mechanism (RQ2-P3) and Churn-Resilient Agent Selection (RQ3)
**Draft — companion methods, written BLIND. Results/Discussion HELD until each track clears its human gate.** **Draft — companion methods. Both tracks have cleared their human gates (RQ2-P3 freeze; RQ3 operator-GO); Results/Discussion are UN-BLINDED and filled from the sealed records only.**
> **Paper-structure note (deliberately left OPEN).** Whether this material ships as a second > **Paper-structure note (deliberately left OPEN).** Whether this material ships as a second
> standalone paper, as extension sections folded into the lead paper > standalone paper, as extension sections folded into the lead paper
@@ -16,19 +16,22 @@
> **offline + deterministic** and its record is **sealed** (`output/sor-rq2p3-confirmatory/…`, > **offline + deterministic** and its record is **sealed** (`output/sor-rq2p3-confirmatory/…`,
> results SHA-256 `5fdcb379d8a2…`). §5/§6 for RQ2-P3 are filled **from that sealed record only**, > results SHA-256 `5fdcb379d8a2…`). §5/§6 for RQ2-P3 are filled **from that sealed record only**,
> the same post-seal discipline the lead paper used. > the same post-seal discipline the lead paper used.
> - **RQ3 companion** — hypotheses, gates, and analysis are **already frozen** in the lead prereg > - **RQ3 companion — RUN + UN-BLINDED.** Hypotheses, gates, and analysis are **frozen** in the lead
> (`sor-consent-prereg.md`, SHA-256 > prereg (`sor-consent-prereg.md`, SHA-256
> `f22331a72e0d0ccf38b787e63acabbe9d666456ec76076787a6d545c3193425b`, §3/§4/§6); the two open > `f22331a72e0d0ccf38b787e63acabbe9d666456ec76076787a6d545c3193425b`, §3/§4/§6); the two open
> `[APPROVAL]` execution params were pinned blind in `docs/rq3-companion-run-brief.md`. **No > `[APPROVAL]` execution params were pinned blind in `docs/rq3-companion-run-brief.md`. The
> confirmatory battery** runs until operator-GO on a live isolated grid (the added-latency DV is > confirmatory battery ran **operator-GO'd on the live isolated docker grid** (3 arms × R=30 ×
> live-only). > C=50 = 4,500 real circuits, `live-docker-e2e`) and its record is **sealed**
> - **All Results / Discussion below are HELD-BLIND placeholders.** The **only** numbers written > (`output/sor-rq3-confirmatory/…`, battery SHA-256 `5b61e461…`, analysis SHA-256 `e09c66ef…`).
> here are already-produced **calibration-gate** values, and every one is labelled *calibration*, > §5/§6 for RQ3 are filled **from that sealed record only**.
> never a confirmatory result. The frozen lead prereg is authoritative and **unedited**. > - **All Results / Discussion below are now UN-BLINDED, filled from the sealed records only.** Both
> tracks have cleared their human gates (RQ2-P3 freeze; RQ3 operator-GO); the **authoritative
> Holm-7** over the frozen size-7 family is computed. The frozen lead prereg is authoritative and
> **unedited**; the lead RQ1/RQ2-P1 findings are not re-litigated.
--- ---
## Abstract *(skeleton — HELD-BLIND; quantitative confirmatory claims withheld until each gate clears)* ## Abstract *(both tracks UN-BLINDED — confirmatory findings folded in from the sealed records)*
The lead study measured a consent-gated, federated, nested-SSH relay instrument and reported two The lead study measured a consent-gated, federated, nested-SSH relay instrument and reported two
honest non-confirmations: no measurable entry↔exit linkability leak (RQ1) and a Holm-significant honest non-confirmations: no measurable entry↔exit linkability leak (RQ1) and a Holm-significant
@@ -44,8 +47,11 @@ dose-response over pool size and skew), asking **two-sided** whether concentrati
whether a **local open-weight agent** path-selector (`qwen2.5:3b`) retains throughput and adds whether a **local open-weight agent** path-selector (`qwen2.5:3b`) retains throughput and adds
tolerable latency under a pinned churn schedule, without leaving a classifiable **rebuild tolerable latency under a pinned churn schedule, without leaving a classifiable **rebuild
fingerprint**. Both tracks are pre-registered, detector-frozen, and calibration-gated before any fingerprint**. Both tracks are pre-registered, detector-frozen, and calibration-gated before any
confirmatory cell. *(Confirmatory findings withheld: RQ2-P3 pending freeze; RQ3 pending confirmatory cell. *(Confirmatory findings, now un-blinded: **RQ2-P3 resolves MIX** —
operator-GO on the live grid.)* shared-pool concentration raises the anonymity set, correcting the lead "shrink" as a
unique-bridge artifact; **RQ3 is a null on both counts** — on this grid every selector
heals ~all churn (no +10 pp agent margin) and the rebuild-timing fingerprint cannot be
excluded at n=30. The authoritative Holm-7 leaves RQ1-P1, RQ2-P1, and RQ2-P3 surviving.)*
--- ---
@@ -239,7 +245,7 @@ confirmatory battery and have both passed on a **dry, synthetic, offline** pass:
--- ---
## 5. Results *(HELD-BLIND — no confirmatory number written)* ## 5. Results *(both tracks UN-BLINDEDfilled from the sealed records only)*
- **RQ2-P3 (H1 / H2 / H3) — RESOLVED: MIX.** From the sealed confirmatory record (9 cells × R=30 × - **RQ2-P3 (H1 / H2 / H3) — RESOLVED: MIX.** From the sealed confirmatory record (9 cells × R=30 ×
C=50, offline + deterministic, S0 = 20260719): C=50, offline + deterministic, S0 = 20260719):
@@ -256,13 +262,38 @@ confirmatory battery and have both passed on a **dry, synthetic, offline** pass:
- Across the sweep, as pool size B rises concentration falls **and** entropy H falls together - Across the sweep, as pool size B rises concentration falls **and** entropy H falls together
(e.g. B=2/α=0: conc ≈ 1.00, H ≈ 2.54; B=8/α=0: conc ≈ 0.51, H ≈ 2.19) — concentration and H move (e.g. B=2/α=0: conc ≈ 1.00, H ≈ 2.54; B=8/α=0: conc ≈ 0.51, H ≈ 2.19) — concentration and H move
**together, positively**: higher concentration ⇒ higher anonymity (mix), not lower (funnel). **together, positively**: higher concentration ⇒ higher anonymity (mix), not lower (funnel).
- **RQ3-P1-perf / RQ3-P1-latency / RQ3-P2 / RQ3-P3.** *HELD — pending operator-GO on the live - **RQ3-P1-perf / RQ3-P1-latency / RQ3-P2 / RQ3-P3 — H0 (honest null).** From the sealed live
isolated grid.* Throughput-retention margin, added-latency (live-only), and rebuild-classifier battery (3 selector arms × R=30 × C=50 = 4,500 real isolated-docker circuits, `measured_from =
AUC over confirmatory cells are **not** computed; the only figures on record are the two green live-docker-e2e`; agent = `qwen2.5:3b` local Ollama; run-level multi-arm bootstrap, 10,000 BCa
**calibration** gates (churn-bites; classifier AUC ≈0.93 sep / ≈0.52 null), labelled calibration. resamples, α = 0.05; results SHA-256 `e09c66ef…`):
- **Holm-7 (companion, authoritative).** *HELD* until all seven confirmatory p-values exist. - **RQ3-P1-perf — FAILS the +10 pp gate.** Throughput-retention margin = retention(agent)
max(static, random) = **0.6 pp**, BCa 95% CI **[1.58 pp, +0.39 pp]**. Every selector heals
~all churn drops (mean retention ≈ 0.99 across arms), so the agent shows **no** ≥ +10 pp gain.
- **RQ3-P1-latency — WITHIN the ≤ 100 ms budget.** Added-latency(agent) = median e2e
latency(agent) median latency(min-latency baseline = random) = **13.5 ms**, BCa 95% CI
**[52.1, +34.9] ms**; CI upper 34.9 ms ≤ 100 ms — the agent is **not** slower than the best
baseline (the perf gate, not latency, is what fails P1).
- **RQ3-P2 — FAILS the ≤ 0.60 ceiling (fingerprint not excluded).** Rebuild-classifier AUC (agent
per-run mean inter-rebuild-gap vs. the pooled baseline selectors) = **0.587**, BCa 95% CI
**[0.458, 0.703]**; CI upper 0.703 **> 0.60**, so a rebuild-timing fingerprint of the agent
**cannot be ruled out** at the pre-registered bar (the test is underpowered at n = 30 runs/arm).
*(Disclosure: the green §3-4 **calibration** gate reported AUC ≈ 0.93 — but that was
**churned-vs-low-churn regime** discrimination on labelled control signals, a **different**
comparison from this confirmatory **agent-vs-baseline-selector** AUC of 0.587; the calibration
validated the instrument and does **not** preview the confirmatory selector value.)*
- **RQ3-P3 (joint) — H0.** P1 fails (perf) **and** P2 fails → the agent selector is **not**
confirmed to help without a fingerprint. Reported as the finding, not spun.
- **Holm-7 (companion, authoritative).** Over the frozen size-7 family {RQ1-P1, RQ1-P2, RQ2-P1,
RQ2-P3, RQ3-P1-perf, RQ3-P1-latency, RQ3-P2}, exact Holm step-down (the RQ2-P3 slot carries the
**mechanism-corrected** primary H1-pooled Spearman p, superseding the lead's degenerate
as-instrumented RQ2-P3): **survivors = RQ1-P1 (rank 1 ×7), RQ2-P1 shrink (rank 2 ×6), RQ2-P3 mix
(rank 3 ×5)** — all adjusted p = 0. **Do not survive:** RQ1-P2 (rank 4 ×4, adj p = 0.365), RQ3-P2
(rank 5 ×3, 0.511), RQ3-P1-perf (rank 6 ×2, 0.511), RQ3-P1-latency (rank 7 ×1, 0.511). This is the
**authoritative** final correction and **supersedes** the lead paper's conservative *partial*
embedding (report-4); both remain valid, the partial never under-corrects, and lead RQ1-P1 /
RQ2-P1 survive regardless.
## 6. Discussion *(HELD-BLIND — placeholders)* ## 6. Discussion *(both tracks UN-BLINDED)*
- **Does a shared bridge funnel or mix? — It mixes.** Both pre-registered two-sided tests resolve - **Does a shared bridge funnel or mix? — It mixes.** Both pre-registered two-sided tests resolve
**positive** (H1 ρ = +0.6244 CI [+0.5941, +0.6545]; H2 β = +0.7052 CI [+0.6195, +0.7903]; H3 **positive** (H1 ρ = +0.6244 CI [+0.5941, +0.6545]; H2 β = +0.7052 CI [+0.6195, +0.7903]; H3
@@ -284,13 +315,39 @@ confirmatory battery and have both passed on a **dry, synthetic, offline** pass:
**quantifies a dose-response that was already visible at calibration**; the pre-committed **quantifies a dose-response that was already visible at calibration**; the pre-committed
hypotheses were nonetheless **two-sided** and that pre-commitment is unchanged. We surface the hypotheses were nonetheless **two-sided** and that pre-commitment is unchanged. We surface the
calibration preview openly so no reader mistakes the confirmatory sign for a post-hoc choice. calibration preview openly so no reader mistakes the confirmatory sign for a post-hoc choice.
- **Does the agent selector help without leaking?** *HELD.* RQ3-P3 confirms only if the agent both - **Does the agent selector help without leaking? — No (H0), on both counts.** RQ3-P3 requires the
clears the +10 pp retention / ≤100 ms latency bar **and** leaves a non-classifiable rebuild agent to clear the +10 pp retention / ≤ 100 ms latency bar **and** leave a non-classifiable
pattern (AUC CI upper ≤ 0.60); either failure is an honest H0 (perf cancelled by teardown rebuild pattern (AUC CI upper ≤ 0.60); it does neither decisively. On **performance**, the honest
overhead, or the rebuild timing is a usable fingerprint). reason is that churn at kp = 30 / steps = 20 is **fully healed by every arm** — static, random,
- **Scope & limitations.** *HELD* — will inherit the lead paper's lab-grid external-validity scope, and the agent all rebuild ~all dropped hops (retention ≈ 0.99), so there is simply **no headroom**
add the RQ2-P3 as-instrumented concentration/mix caveat, and carry the RQ3 agent-reproducibility for an adaptive selector to win the +10 pp margin (margin 0.6 pp, CI [1.58, +0.39] pp). The
caveat (cache-replay, not cross-hardware re-execution) with equal prominence. agent is not *worse* — added-latency is within budget (13.5 ms, CI upper 34.9 ms ≤ 100 ms) — it
is merely **not better**, because the baseline is already at the retention ceiling on this grid.
On **anonymity**, the rebuild-timing classifier reaches AUC 0.587 with CI upper 0.703 > 0.60, so
the pre-registered bar to certify "no usable fingerprint" is **not met**: at n = 30 runs/arm the
test is **underpowered** to exclude a small rebuild-timing signal, and we report that limitation
rather than a false all-clear. The honest reading: **on this lab grid the local open-weight agent
selector neither beats the baselines nor is demonstrably fingerprint-free** — a null on both P1
and P2, exactly the outcome the design pre-committed to publish with equal prominence.
- **Authoritative multiplicity (Holm-7).** With all seven pre-registered p-values now in hand, the
companion's exact Holm-7 supersedes the lead paper's conservative partial embedding (operator
decision D3). Three hypotheses survive: **RQ1-P1** (below-chance linkability AUC — no usable
entry↔exit leak), **RQ2-P1** (federation *shrinks* the anonymity set, the lead headline null),
and **RQ2-P3** (shared-pool concentration *mixes* — the mechanism correction). The four that do
not survive are RQ1-P2 (padding efficacy), and all three RQ3 tests — consistent with the RQ3 H0
above. Note the RQ2-P3 slot now carries the **mechanism-corrected** primary statistic (H1-pooled
Spearman, adj p = 0) rather than the lead's degenerate as-instrumented RQ2-P3 (adj p = 1): the
companion's frozen-detector method both *caught* the unique-bridge artifact and *promotes* the
corrected mechanism finding into the surviving family. Lead RQ1-P1 and RQ2-P1 survive regardless.
- **Scope & limitations.** Both findings are scoped to the lab grid (1-house / bridge-off control,
2 phones + laptop, self-generated fixture traffic) and inherit the lead paper's external-validity
caveats. Specific to this companion: (i) the RQ2-P3 mix is an **as-instrumented** concentration
effect on the ratified exit-signature posterior, not an internet-scale claim; (ii) the RQ3 nulls
are **grid-bound** — the perf null follows from a baseline retention ceiling under the pinned
churn, and the P2 non-exclusion is an n = 30 **power** limitation, not a proof of a fingerprint;
(iii) the agent arm carries the accepted **reproducibility caveat** (Ollama temp-0 is reproducible
via the committed decision-log + (seed, state-hash) cache replay, *not* via cross-hardware model
re-execution), stated with equal prominence to the RQ1 timing caveat.
--- ---