paper: fill RQ3 half of combined companion (post-seal, both tracks un-blinded)

Replace the HELD-BLIND RQ3 Results/Discussion + authoritative-Holm-7
placeholders in docs/stage-07-companion-methods.md with real numbers read
from the sealed record only (rq3-confirmatory-analysis.json e09c66ef…):

  RQ3-P1-perf   -0.6pp  CI[-1.58,+0.39]pp   -> H0 (no +10pp agent margin)
  RQ3-P1-latency -13.5ms CI[-52.1,+34.9]ms  -> within 100ms budget
  RQ3-P2 AUC     0.587  CI[0.458,0.703]      -> fingerprint NOT excluded
  RQ3-P3 = H0 (perf fails AND P2 fails)

Authoritative Holm-7 (family size 7, supersedes lead partial embedding, D3):
survivors RQ1-P1 / RQ2-P1 / RQ2-P3; non-survivors RQ1-P2 / RQ3-P2 /
RQ3-P1-perf / RQ3-P1-latency. RQ2-P3 slot carries the mechanism-corrected
primary p, not the lead degenerate p=1.

Calibration disclosure carried: the 0.93 RQ3 calibration AUC is regime
discrimination (churn kp30 vs kp5), distinct from the 0.587 confirmatory
selector-vs-selector fingerprint — no HARKing, detectors frozen.

RQ2-P3 half not re-litigated (kept as filled at e0d865d). Nulls reported
openly as results. Frozen lead prereg f22331a72e… and RQ2-P3 prereg
8db4e8a7… both intact. $0/offline, worktree-only.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
This commit is contained in:
leetcrypt
2026-07-22 08:53:50 -07:00
parent 47faee3a92
commit 1b6449b827
2 changed files with 84 additions and 26 deletions
+83 -26
View File
@@ -1,6 +1,6 @@
# Companion Methods (BLIND scaffold): The Unique-Bridge / Mix Mechanism (RQ2-P3) and Churn-Resilient Agent Selection (RQ3)
**Draft — companion methods, written BLIND. Results/Discussion HELD until each track clears its human gate.**
**Draft — companion methods. Both tracks have cleared their human gates (RQ2-P3 freeze; RQ3 operator-GO); Results/Discussion are UN-BLINDED and filled from the sealed records only.**
> **Paper-structure note (deliberately left OPEN).** Whether this material ships as a second
> standalone paper, as extension sections folded into the lead paper
@@ -16,19 +16,22 @@
> **offline + deterministic** and its record is **sealed** (`output/sor-rq2p3-confirmatory/…`,
> results SHA-256 `5fdcb379d8a2…`). §5/§6 for RQ2-P3 are filled **from that sealed record only**,
> the same post-seal discipline the lead paper used.
> - **RQ3 companion** — hypotheses, gates, and analysis are **already frozen** in the lead prereg
> (`sor-consent-prereg.md`, SHA-256
> - **RQ3 companion — RUN + UN-BLINDED.** Hypotheses, gates, and analysis are **frozen** in the lead
> prereg (`sor-consent-prereg.md`, SHA-256
> `f22331a72e0d0ccf38b787e63acabbe9d666456ec76076787a6d545c3193425b`, §3/§4/§6); the two open
> `[APPROVAL]` execution params were pinned blind in `docs/rq3-companion-run-brief.md`. **No
> confirmatory battery** runs until operator-GO on a live isolated grid (the added-latency DV is
> live-only).
> - **All Results / Discussion below are HELD-BLIND placeholders.** The **only** numbers written
> here are already-produced **calibration-gate** values, and every one is labelled *calibration*,
> never a confirmatory result. The frozen lead prereg is authoritative and **unedited**.
> `[APPROVAL]` execution params were pinned blind in `docs/rq3-companion-run-brief.md`. The
> confirmatory battery ran **operator-GO'd on the live isolated docker grid** (3 arms × R=30 ×
> C=50 = 4,500 real circuits, `live-docker-e2e`) and its record is **sealed**
> (`output/sor-rq3-confirmatory/…`, battery SHA-256 `5b61e461…`, analysis SHA-256 `e09c66ef…`).
> §5/§6 for RQ3 are filled **from that sealed record only**.
> - **All Results / Discussion below are now UN-BLINDED, filled from the sealed records only.** Both
> tracks have cleared their human gates (RQ2-P3 freeze; RQ3 operator-GO); the **authoritative
> Holm-7** over the frozen size-7 family is computed. The frozen lead prereg is authoritative and
> **unedited**; the lead RQ1/RQ2-P1 findings are not re-litigated.
---
## Abstract *(skeleton — HELD-BLIND; quantitative confirmatory claims withheld until each gate clears)*
## Abstract *(both tracks UN-BLINDED — confirmatory findings folded in from the sealed records)*
The lead study measured a consent-gated, federated, nested-SSH relay instrument and reported two
honest non-confirmations: no measurable entry↔exit linkability leak (RQ1) and a Holm-significant
@@ -44,8 +47,11 @@ dose-response over pool size and skew), asking **two-sided** whether concentrati
whether a **local open-weight agent** path-selector (`qwen2.5:3b`) retains throughput and adds
tolerable latency under a pinned churn schedule, without leaving a classifiable **rebuild
fingerprint**. Both tracks are pre-registered, detector-frozen, and calibration-gated before any
confirmatory cell. *(Confirmatory findings withheld: RQ2-P3 pending freeze; RQ3 pending
operator-GO on the live grid.)*
confirmatory cell. *(Confirmatory findings, now un-blinded: **RQ2-P3 resolves MIX** —
shared-pool concentration raises the anonymity set, correcting the lead "shrink" as a
unique-bridge artifact; **RQ3 is a null on both counts** — on this grid every selector
heals ~all churn (no +10 pp agent margin) and the rebuild-timing fingerprint cannot be
excluded at n=30. The authoritative Holm-7 leaves RQ1-P1, RQ2-P1, and RQ2-P3 surviving.)*
---
@@ -239,7 +245,7 @@ confirmatory battery and have both passed on a **dry, synthetic, offline** pass:
---
## 5. Results *(HELD-BLIND — no confirmatory number written)*
## 5. Results *(both tracks UN-BLINDEDfilled from the sealed records only)*
- **RQ2-P3 (H1 / H2 / H3) — RESOLVED: MIX.** From the sealed confirmatory record (9 cells × R=30 ×
C=50, offline + deterministic, S0 = 20260719):
@@ -256,13 +262,38 @@ confirmatory battery and have both passed on a **dry, synthetic, offline** pass:
- Across the sweep, as pool size B rises concentration falls **and** entropy H falls together
(e.g. B=2/α=0: conc ≈ 1.00, H ≈ 2.54; B=8/α=0: conc ≈ 0.51, H ≈ 2.19) — concentration and H move
**together, positively**: higher concentration ⇒ higher anonymity (mix), not lower (funnel).
- **RQ3-P1-perf / RQ3-P1-latency / RQ3-P2 / RQ3-P3.** *HELD — pending operator-GO on the live
isolated grid.* Throughput-retention margin, added-latency (live-only), and rebuild-classifier
AUC over confirmatory cells are **not** computed; the only figures on record are the two green
**calibration** gates (churn-bites; classifier AUC ≈0.93 sep / ≈0.52 null), labelled calibration.
- **Holm-7 (companion, authoritative).** *HELD* until all seven confirmatory p-values exist.
- **RQ3-P1-perf / RQ3-P1-latency / RQ3-P2 / RQ3-P3 — H0 (honest null).** From the sealed live
battery (3 selector arms × R=30 × C=50 = 4,500 real isolated-docker circuits, `measured_from =
live-docker-e2e`; agent = `qwen2.5:3b` local Ollama; run-level multi-arm bootstrap, 10,000 BCa
resamples, α = 0.05; results SHA-256 `e09c66ef…`):
- **RQ3-P1-perf — FAILS the +10 pp gate.** Throughput-retention margin = retention(agent)
max(static, random) = **0.6 pp**, BCa 95% CI **[1.58 pp, +0.39 pp]**. Every selector heals
~all churn drops (mean retention ≈ 0.99 across arms), so the agent shows **no** ≥ +10 pp gain.
- **RQ3-P1-latency — WITHIN the ≤ 100 ms budget.** Added-latency(agent) = median e2e
latency(agent) median latency(min-latency baseline = random) = **13.5 ms**, BCa 95% CI
**[52.1, +34.9] ms**; CI upper 34.9 ms ≤ 100 ms — the agent is **not** slower than the best
baseline (the perf gate, not latency, is what fails P1).
- **RQ3-P2 — FAILS the ≤ 0.60 ceiling (fingerprint not excluded).** Rebuild-classifier AUC (agent
per-run mean inter-rebuild-gap vs. the pooled baseline selectors) = **0.587**, BCa 95% CI
**[0.458, 0.703]**; CI upper 0.703 **> 0.60**, so a rebuild-timing fingerprint of the agent
**cannot be ruled out** at the pre-registered bar (the test is underpowered at n = 30 runs/arm).
*(Disclosure: the green §3-4 **calibration** gate reported AUC ≈ 0.93 — but that was
**churned-vs-low-churn regime** discrimination on labelled control signals, a **different**
comparison from this confirmatory **agent-vs-baseline-selector** AUC of 0.587; the calibration
validated the instrument and does **not** preview the confirmatory selector value.)*
- **RQ3-P3 (joint) — H0.** P1 fails (perf) **and** P2 fails → the agent selector is **not**
confirmed to help without a fingerprint. Reported as the finding, not spun.
- **Holm-7 (companion, authoritative).** Over the frozen size-7 family {RQ1-P1, RQ1-P2, RQ2-P1,
RQ2-P3, RQ3-P1-perf, RQ3-P1-latency, RQ3-P2}, exact Holm step-down (the RQ2-P3 slot carries the
**mechanism-corrected** primary H1-pooled Spearman p, superseding the lead's degenerate
as-instrumented RQ2-P3): **survivors = RQ1-P1 (rank 1 ×7), RQ2-P1 shrink (rank 2 ×6), RQ2-P3 mix
(rank 3 ×5)** — all adjusted p = 0. **Do not survive:** RQ1-P2 (rank 4 ×4, adj p = 0.365), RQ3-P2
(rank 5 ×3, 0.511), RQ3-P1-perf (rank 6 ×2, 0.511), RQ3-P1-latency (rank 7 ×1, 0.511). This is the
**authoritative** final correction and **supersedes** the lead paper's conservative *partial*
embedding (report-4); both remain valid, the partial never under-corrects, and lead RQ1-P1 /
RQ2-P1 survive regardless.
## 6. Discussion *(HELD-BLIND — placeholders)*
## 6. Discussion *(both tracks UN-BLINDED)*
- **Does a shared bridge funnel or mix? — It mixes.** Both pre-registered two-sided tests resolve
**positive** (H1 ρ = +0.6244 CI [+0.5941, +0.6545]; H2 β = +0.7052 CI [+0.6195, +0.7903]; H3
@@ -284,13 +315,39 @@ confirmatory battery and have both passed on a **dry, synthetic, offline** pass:
**quantifies a dose-response that was already visible at calibration**; the pre-committed
hypotheses were nonetheless **two-sided** and that pre-commitment is unchanged. We surface the
calibration preview openly so no reader mistakes the confirmatory sign for a post-hoc choice.
- **Does the agent selector help without leaking?** *HELD.* RQ3-P3 confirms only if the agent both
clears the +10 pp retention / ≤100 ms latency bar **and** leaves a non-classifiable rebuild
pattern (AUC CI upper ≤ 0.60); either failure is an honest H0 (perf cancelled by teardown
overhead, or the rebuild timing is a usable fingerprint).
- **Scope & limitations.** *HELD* — will inherit the lead paper's lab-grid external-validity scope,
add the RQ2-P3 as-instrumented concentration/mix caveat, and carry the RQ3 agent-reproducibility
caveat (cache-replay, not cross-hardware re-execution) with equal prominence.
- **Does the agent selector help without leaking? — No (H0), on both counts.** RQ3-P3 requires the
agent to clear the +10 pp retention / ≤ 100 ms latency bar **and** leave a non-classifiable
rebuild pattern (AUC CI upper ≤ 0.60); it does neither decisively. On **performance**, the honest
reason is that churn at kp = 30 / steps = 20 is **fully healed by every arm** — static, random,
and the agent all rebuild ~all dropped hops (retention ≈ 0.99), so there is simply **no headroom**
for an adaptive selector to win the +10 pp margin (margin 0.6 pp, CI [1.58, +0.39] pp). The
agent is not *worse* — added-latency is within budget (13.5 ms, CI upper 34.9 ms ≤ 100 ms) — it
is merely **not better**, because the baseline is already at the retention ceiling on this grid.
On **anonymity**, the rebuild-timing classifier reaches AUC 0.587 with CI upper 0.703 > 0.60, so
the pre-registered bar to certify "no usable fingerprint" is **not met**: at n = 30 runs/arm the
test is **underpowered** to exclude a small rebuild-timing signal, and we report that limitation
rather than a false all-clear. The honest reading: **on this lab grid the local open-weight agent
selector neither beats the baselines nor is demonstrably fingerprint-free** — a null on both P1
and P2, exactly the outcome the design pre-committed to publish with equal prominence.
- **Authoritative multiplicity (Holm-7).** With all seven pre-registered p-values now in hand, the
companion's exact Holm-7 supersedes the lead paper's conservative partial embedding (operator
decision D3). Three hypotheses survive: **RQ1-P1** (below-chance linkability AUC — no usable
entry↔exit leak), **RQ2-P1** (federation *shrinks* the anonymity set, the lead headline null),
and **RQ2-P3** (shared-pool concentration *mixes* — the mechanism correction). The four that do
not survive are RQ1-P2 (padding efficacy), and all three RQ3 tests — consistent with the RQ3 H0
above. Note the RQ2-P3 slot now carries the **mechanism-corrected** primary statistic (H1-pooled
Spearman, adj p = 0) rather than the lead's degenerate as-instrumented RQ2-P3 (adj p = 1): the
companion's frozen-detector method both *caught* the unique-bridge artifact and *promotes* the
corrected mechanism finding into the surviving family. Lead RQ1-P1 and RQ2-P1 survive regardless.
- **Scope & limitations.** Both findings are scoped to the lab grid (1-house / bridge-off control,
2 phones + laptop, self-generated fixture traffic) and inherit the lead paper's external-validity
caveats. Specific to this companion: (i) the RQ2-P3 mix is an **as-instrumented** concentration
effect on the ratified exit-signature posterior, not an internet-scale claim; (ii) the RQ3 nulls
are **grid-bound** — the perf null follows from a baseline retention ceiling under the pinned
churn, and the P2 non-exclusion is an n = 30 **power** limitation, not a proof of a fingerprint;
(iii) the agent arm carries the accepted **reproducibility caveat** (Ollama temp-0 is reproducible
via the committed decision-log + (seed, state-hash) cache replay, *not* via cross-hardware model
re-execution), stated with equal prominence to the RQ1 timing caveat.
---