paper: fill RQ3 half of combined companion (post-seal, both tracks un-blinded)
Replace the HELD-BLIND RQ3 Results/Discussion + authoritative-Holm-7
placeholders in docs/stage-07-companion-methods.md with real numbers read
from the sealed record only (rq3-confirmatory-analysis.json e09c66ef…):
RQ3-P1-perf -0.6pp CI[-1.58,+0.39]pp -> H0 (no +10pp agent margin)
RQ3-P1-latency -13.5ms CI[-52.1,+34.9]ms -> within 100ms budget
RQ3-P2 AUC 0.587 CI[0.458,0.703] -> fingerprint NOT excluded
RQ3-P3 = H0 (perf fails AND P2 fails)
Authoritative Holm-7 (family size 7, supersedes lead partial embedding, D3):
survivors RQ1-P1 / RQ2-P1 / RQ2-P3; non-survivors RQ1-P2 / RQ3-P2 /
RQ3-P1-perf / RQ3-P1-latency. RQ2-P3 slot carries the mechanism-corrected
primary p, not the lead degenerate p=1.
Calibration disclosure carried: the 0.93 RQ3 calibration AUC is regime
discrimination (churn kp30 vs kp5), distinct from the 0.587 confirmatory
selector-vs-selector fingerprint — no HARKing, detectors frozen.
RQ2-P3 half not re-litigated (kept as filled at e0d865d). Nulls reported
openly as results. Frozen lead prereg f22331a72e… and RQ2-P3 prereg
8db4e8a7… both intact. $0/offline, worktree-only.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
This commit is contained in:
@@ -1,6 +1,6 @@
|
||||
# Companion Methods (BLIND scaffold): The Unique-Bridge / Mix Mechanism (RQ2-P3) and Churn-Resilient Agent Selection (RQ3)
|
||||
|
||||
**Draft — companion methods, written BLIND. Results/Discussion HELD until each track clears its human gate.**
|
||||
**Draft — companion methods. Both tracks have cleared their human gates (RQ2-P3 freeze; RQ3 operator-GO); Results/Discussion are UN-BLINDED and filled from the sealed records only.**
|
||||
|
||||
> **Paper-structure note (deliberately left OPEN).** Whether this material ships as a second
|
||||
> standalone paper, as extension sections folded into the lead paper
|
||||
@@ -16,19 +16,22 @@
|
||||
> **offline + deterministic** and its record is **sealed** (`output/sor-rq2p3-confirmatory/…`,
|
||||
> results SHA-256 `5fdcb379d8a2…`). §5/§6 for RQ2-P3 are filled **from that sealed record only**,
|
||||
> the same post-seal discipline the lead paper used.
|
||||
> - **RQ3 companion** — hypotheses, gates, and analysis are **already frozen** in the lead prereg
|
||||
> (`sor-consent-prereg.md`, SHA-256
|
||||
> - **RQ3 companion — RUN + UN-BLINDED.** Hypotheses, gates, and analysis are **frozen** in the lead
|
||||
> prereg (`sor-consent-prereg.md`, SHA-256
|
||||
> `f22331a72e0d0ccf38b787e63acabbe9d666456ec76076787a6d545c3193425b`, §3/§4/§6); the two open
|
||||
> `[APPROVAL]` execution params were pinned blind in `docs/rq3-companion-run-brief.md`. **No
|
||||
> confirmatory battery** runs until operator-GO on a live isolated grid (the added-latency DV is
|
||||
> live-only).
|
||||
> - **All Results / Discussion below are HELD-BLIND placeholders.** The **only** numbers written
|
||||
> here are already-produced **calibration-gate** values, and every one is labelled *calibration*,
|
||||
> never a confirmatory result. The frozen lead prereg is authoritative and **unedited**.
|
||||
> `[APPROVAL]` execution params were pinned blind in `docs/rq3-companion-run-brief.md`. The
|
||||
> confirmatory battery ran **operator-GO'd on the live isolated docker grid** (3 arms × R=30 ×
|
||||
> C=50 = 4,500 real circuits, `live-docker-e2e`) and its record is **sealed**
|
||||
> (`output/sor-rq3-confirmatory/…`, battery SHA-256 `5b61e461…`, analysis SHA-256 `e09c66ef…`).
|
||||
> §5/§6 for RQ3 are filled **from that sealed record only**.
|
||||
> - **All Results / Discussion below are now UN-BLINDED, filled from the sealed records only.** Both
|
||||
> tracks have cleared their human gates (RQ2-P3 freeze; RQ3 operator-GO); the **authoritative
|
||||
> Holm-7** over the frozen size-7 family is computed. The frozen lead prereg is authoritative and
|
||||
> **unedited**; the lead RQ1/RQ2-P1 findings are not re-litigated.
|
||||
|
||||
---
|
||||
|
||||
## Abstract *(skeleton — HELD-BLIND; quantitative confirmatory claims withheld until each gate clears)*
|
||||
## Abstract *(both tracks UN-BLINDED — confirmatory findings folded in from the sealed records)*
|
||||
|
||||
The lead study measured a consent-gated, federated, nested-SSH relay instrument and reported two
|
||||
honest non-confirmations: no measurable entry↔exit linkability leak (RQ1) and a Holm-significant
|
||||
@@ -44,8 +47,11 @@ dose-response over pool size and skew), asking **two-sided** whether concentrati
|
||||
whether a **local open-weight agent** path-selector (`qwen2.5:3b`) retains throughput and adds
|
||||
tolerable latency under a pinned churn schedule, without leaving a classifiable **rebuild
|
||||
fingerprint**. Both tracks are pre-registered, detector-frozen, and calibration-gated before any
|
||||
confirmatory cell. *(Confirmatory findings withheld: RQ2-P3 pending freeze; RQ3 pending
|
||||
operator-GO on the live grid.)*
|
||||
confirmatory cell. *(Confirmatory findings, now un-blinded: **RQ2-P3 resolves MIX** —
|
||||
shared-pool concentration raises the anonymity set, correcting the lead "shrink" as a
|
||||
unique-bridge artifact; **RQ3 is a null on both counts** — on this grid every selector
|
||||
heals ~all churn (no +10 pp agent margin) and the rebuild-timing fingerprint cannot be
|
||||
excluded at n=30. The authoritative Holm-7 leaves RQ1-P1, RQ2-P1, and RQ2-P3 surviving.)*
|
||||
|
||||
---
|
||||
|
||||
@@ -239,7 +245,7 @@ confirmatory battery and have both passed on a **dry, synthetic, offline** pass:
|
||||
|
||||
---
|
||||
|
||||
## 5. Results *(HELD-BLIND — no confirmatory number written)*
|
||||
## 5. Results *(both tracks UN-BLINDED — filled from the sealed records only)*
|
||||
|
||||
- **RQ2-P3′ (H1 / H2 / H3) — RESOLVED: MIX.** From the sealed confirmatory record (9 cells × R=30 ×
|
||||
C=50, offline + deterministic, S0 = 20260719):
|
||||
@@ -256,13 +262,38 @@ confirmatory battery and have both passed on a **dry, synthetic, offline** pass:
|
||||
- Across the sweep, as pool size B rises concentration falls **and** entropy H falls together
|
||||
(e.g. B=2/α=0: conc ≈ 1.00, H ≈ 2.54; B=8/α=0: conc ≈ 0.51, H ≈ 2.19) — concentration and H move
|
||||
**together, positively**: higher concentration ⇒ higher anonymity (mix), not lower (funnel).
|
||||
- **RQ3-P1-perf / RQ3-P1-latency / RQ3-P2 / RQ3-P3.** *HELD — pending operator-GO on the live
|
||||
isolated grid.* Throughput-retention margin, added-latency (live-only), and rebuild-classifier
|
||||
AUC over confirmatory cells are **not** computed; the only figures on record are the two green
|
||||
**calibration** gates (churn-bites; classifier AUC ≈0.93 sep / ≈0.52 null), labelled calibration.
|
||||
- **Holm-7 (companion, authoritative).** *HELD* until all seven confirmatory p-values exist.
|
||||
- **RQ3-P1-perf / RQ3-P1-latency / RQ3-P2 / RQ3-P3 — H0 (honest null).** From the sealed live
|
||||
battery (3 selector arms × R=30 × C=50 = 4,500 real isolated-docker circuits, `measured_from =
|
||||
live-docker-e2e`; agent = `qwen2.5:3b` local Ollama; run-level multi-arm bootstrap, 10,000 BCa
|
||||
resamples, α = 0.05; results SHA-256 `e09c66ef…`):
|
||||
- **RQ3-P1-perf — FAILS the +10 pp gate.** Throughput-retention margin = retention(agent) −
|
||||
max(static, random) = **−0.6 pp**, BCa 95% CI **[−1.58 pp, +0.39 pp]**. Every selector heals
|
||||
~all churn drops (mean retention ≈ 0.99 across arms), so the agent shows **no** ≥ +10 pp gain.
|
||||
- **RQ3-P1-latency — WITHIN the ≤ 100 ms budget.** Added-latency(agent) = median e2e
|
||||
latency(agent) − median latency(min-latency baseline = random) = **−13.5 ms**, BCa 95% CI
|
||||
**[−52.1, +34.9] ms**; CI upper 34.9 ms ≤ 100 ms — the agent is **not** slower than the best
|
||||
baseline (the perf gate, not latency, is what fails P1).
|
||||
- **RQ3-P2 — FAILS the ≤ 0.60 ceiling (fingerprint not excluded).** Rebuild-classifier AUC (agent
|
||||
per-run mean inter-rebuild-gap vs. the pooled baseline selectors) = **0.587**, BCa 95% CI
|
||||
**[0.458, 0.703]**; CI upper 0.703 **> 0.60**, so a rebuild-timing fingerprint of the agent
|
||||
**cannot be ruled out** at the pre-registered bar (the test is underpowered at n = 30 runs/arm).
|
||||
*(Disclosure: the green §3-4 **calibration** gate reported AUC ≈ 0.93 — but that was
|
||||
**churned-vs-low-churn regime** discrimination on labelled control signals, a **different**
|
||||
comparison from this confirmatory **agent-vs-baseline-selector** AUC of 0.587; the calibration
|
||||
validated the instrument and does **not** preview the confirmatory selector value.)*
|
||||
- **RQ3-P3 (joint) — H0.** P1 fails (perf) **and** P2 fails → the agent selector is **not**
|
||||
confirmed to help without a fingerprint. Reported as the finding, not spun.
|
||||
- **Holm-7 (companion, authoritative).** Over the frozen size-7 family {RQ1-P1, RQ1-P2, RQ2-P1,
|
||||
RQ2-P3, RQ3-P1-perf, RQ3-P1-latency, RQ3-P2}, exact Holm step-down (the RQ2-P3 slot carries the
|
||||
**mechanism-corrected** primary H1-pooled Spearman p, superseding the lead's degenerate
|
||||
as-instrumented RQ2-P3): **survivors = RQ1-P1 (rank 1 ×7), RQ2-P1 shrink (rank 2 ×6), RQ2-P3 mix
|
||||
(rank 3 ×5)** — all adjusted p = 0. **Do not survive:** RQ1-P2 (rank 4 ×4, adj p = 0.365), RQ3-P2
|
||||
(rank 5 ×3, 0.511), RQ3-P1-perf (rank 6 ×2, 0.511), RQ3-P1-latency (rank 7 ×1, 0.511). This is the
|
||||
**authoritative** final correction and **supersedes** the lead paper's conservative *partial*
|
||||
embedding (report-4); both remain valid, the partial never under-corrects, and lead RQ1-P1 /
|
||||
RQ2-P1 survive regardless.
|
||||
|
||||
## 6. Discussion *(HELD-BLIND — placeholders)*
|
||||
## 6. Discussion *(both tracks UN-BLINDED)*
|
||||
|
||||
- **Does a shared bridge funnel or mix? — It mixes.** Both pre-registered two-sided tests resolve
|
||||
**positive** (H1 ρ = +0.6244 CI [+0.5941, +0.6545]; H2 β = +0.7052 CI [+0.6195, +0.7903]; H3
|
||||
@@ -284,13 +315,39 @@ confirmatory battery and have both passed on a **dry, synthetic, offline** pass:
|
||||
**quantifies a dose-response that was already visible at calibration**; the pre-committed
|
||||
hypotheses were nonetheless **two-sided** and that pre-commitment is unchanged. We surface the
|
||||
calibration preview openly so no reader mistakes the confirmatory sign for a post-hoc choice.
|
||||
- **Does the agent selector help without leaking?** *HELD.* RQ3-P3 confirms only if the agent both
|
||||
clears the +10 pp retention / ≤100 ms latency bar **and** leaves a non-classifiable rebuild
|
||||
pattern (AUC CI upper ≤ 0.60); either failure is an honest H0 (perf cancelled by teardown
|
||||
overhead, or the rebuild timing is a usable fingerprint).
|
||||
- **Scope & limitations.** *HELD* — will inherit the lead paper's lab-grid external-validity scope,
|
||||
add the RQ2-P3 as-instrumented concentration/mix caveat, and carry the RQ3 agent-reproducibility
|
||||
caveat (cache-replay, not cross-hardware re-execution) with equal prominence.
|
||||
- **Does the agent selector help without leaking? — No (H0), on both counts.** RQ3-P3 requires the
|
||||
agent to clear the +10 pp retention / ≤ 100 ms latency bar **and** leave a non-classifiable
|
||||
rebuild pattern (AUC CI upper ≤ 0.60); it does neither decisively. On **performance**, the honest
|
||||
reason is that churn at kp = 30 / steps = 20 is **fully healed by every arm** — static, random,
|
||||
and the agent all rebuild ~all dropped hops (retention ≈ 0.99), so there is simply **no headroom**
|
||||
for an adaptive selector to win the +10 pp margin (margin −0.6 pp, CI [−1.58, +0.39] pp). The
|
||||
agent is not *worse* — added-latency is within budget (−13.5 ms, CI upper 34.9 ms ≤ 100 ms) — it
|
||||
is merely **not better**, because the baseline is already at the retention ceiling on this grid.
|
||||
On **anonymity**, the rebuild-timing classifier reaches AUC 0.587 with CI upper 0.703 > 0.60, so
|
||||
the pre-registered bar to certify "no usable fingerprint" is **not met**: at n = 30 runs/arm the
|
||||
test is **underpowered** to exclude a small rebuild-timing signal, and we report that limitation
|
||||
rather than a false all-clear. The honest reading: **on this lab grid the local open-weight agent
|
||||
selector neither beats the baselines nor is demonstrably fingerprint-free** — a null on both P1
|
||||
and P2, exactly the outcome the design pre-committed to publish with equal prominence.
|
||||
- **Authoritative multiplicity (Holm-7).** With all seven pre-registered p-values now in hand, the
|
||||
companion's exact Holm-7 supersedes the lead paper's conservative partial embedding (operator
|
||||
decision D3). Three hypotheses survive: **RQ1-P1** (below-chance linkability AUC — no usable
|
||||
entry↔exit leak), **RQ2-P1** (federation *shrinks* the anonymity set, the lead headline null),
|
||||
and **RQ2-P3** (shared-pool concentration *mixes* — the mechanism correction). The four that do
|
||||
not survive are RQ1-P2 (padding efficacy), and all three RQ3 tests — consistent with the RQ3 H0
|
||||
above. Note the RQ2-P3 slot now carries the **mechanism-corrected** primary statistic (H1-pooled
|
||||
Spearman, adj p = 0) rather than the lead's degenerate as-instrumented RQ2-P3 (adj p = 1): the
|
||||
companion's frozen-detector method both *caught* the unique-bridge artifact and *promotes* the
|
||||
corrected mechanism finding into the surviving family. Lead RQ1-P1 and RQ2-P1 survive regardless.
|
||||
- **Scope & limitations.** Both findings are scoped to the lab grid (1-house / bridge-off control,
|
||||
2 phones + laptop, self-generated fixture traffic) and inherit the lead paper's external-validity
|
||||
caveats. Specific to this companion: (i) the RQ2-P3 mix is an **as-instrumented** concentration
|
||||
effect on the ratified exit-signature posterior, not an internet-scale claim; (ii) the RQ3 nulls
|
||||
are **grid-bound** — the perf null follows from a baseline retention ceiling under the pinned
|
||||
churn, and the P2 non-exclusion is an n = 30 **power** limitation, not a proof of a fingerprint;
|
||||
(iii) the agent arm carries the accepted **reproducibility caveat** (Ollama temp-0 is reproducible
|
||||
via the committed decision-log + (seed, state-hash) cache replay, *not* via cross-hardware model
|
||||
re-execution), stated with equal prominence to the RQ1 timing caveat.
|
||||
|
||||
---
|
||||
|
||||
|
||||
Reference in New Issue
Block a user