Files
hack-house/docs/stage-07-companion-methods.md
T
leetcrypt 50b1701e6a docs: QA prose fix — attribute lead RQ1-P1/RQ2-P1 raw p to the lead paper
Bounded companion-methods QA (blinding / citation-grounding / FREEZE-PENDING all
clean). Prose-only clarification, no number changed: mark the stray raw p ~ 0 as
the lead paper's already-published RQ1-P1/RQ2-P1 result so it cannot be misread as
a held companion figure. HOLDING — awaiting operator freeze (RQ2-P3 §10) / GO
(RQ3 live grid).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-07-21 20:24:05 -07:00

19 KiB
Raw Blame History

Companion Methods (BLIND scaffold): The Unique-Bridge / Mix Mechanism (RQ2-P3) and Churn-Resilient Agent Selection (RQ3)

Draft — companion methods, written BLIND. Results/Discussion HELD until each track clears its human gate.

Paper-structure note (deliberately left OPEN). Whether this material ships as a second standalone paper, as extension sections folded into the lead paper (docs/stage-07-paper-draft.md), or as a short mechanism note is an operator editorial decision and is not pre-committed here. The two methods tracks below are therefore written as self-contained sections that can be lifted into either structure.

Blinding & gating status (binding).

  • RQ2-P3 mechanism study — the prereg (docs/rq2p3-mechanism-prereg.md, own slug sor-consent-rq2p3) is NOT YET FROZEN (§10 empty). Its methods section below is marked FREEZE-PENDING; no confirmatory cell runs until the operator records the SHA-256 in §10 and signs off.
  • RQ3 companion — hypotheses, gates, and analysis are already frozen in the lead prereg (sor-consent-prereg.md, SHA-256 f22331a72e0d0ccf38b787e63acabbe9d666456ec76076787a6d545c3193425b, §3/§4/§6); the two open [APPROVAL] execution params were pinned blind in docs/rq3-companion-run-brief.md. No confirmatory battery runs until operator-GO on a live isolated grid (the added-latency DV is live-only).
  • All Results / Discussion below are HELD-BLIND placeholders. The only numbers written here are already-produced calibration-gate values, and every one is labelled calibration, never a confirmatory result. The frozen lead prereg is authoritative and unedited.

Abstract (skeleton — HELD-BLIND; quantitative confirmatory claims withheld until each gate clears)

The lead study measured a consent-gated, federated, nested-SSH relay instrument and reported two honest non-confirmations: no measurable entry↔exit linkability leak (RQ1) and a Holm-significant shrink of the per-circuit anonymity set under federation (RQ2-P1). This companion pursues the two questions the lead paper could not close. First (RQ2-P3, a mechanism study): the lead "shrink" may be an instrument artifact — the bridge-federated topology assigns a fresh willing bridge per circuit seed, so every adversary-observable exit signature is unique, every anonymity set collapses to size 1, and entropy is driven to ≈0 by construction rather than by funnelling. We re-instrument the willing-bridge layer as a finite shared pool with skewed willingness and treat bridge concentration as a manipulated independent variable (a 3×3 dose-response over pool size and skew), asking two-sided whether concentration reduces (funnel) or raises (mix) the anonymity set. Second (RQ3, churn resilience): we measure whether a local open-weight agent path-selector (qwen2.5:3b) retains throughput and adds tolerable latency under a pinned churn schedule, without leaving a classifiable rebuild fingerprint. Both tracks are pre-registered, detector-frozen, and calibration-gated before any confirmatory cell. (Confirmatory findings withheld: RQ2-P3 pending freeze; RQ3 pending operator-GO on the live grid.)


1. Introduction (deltas beyond the lead paper)

The lead paper (G4 + RQ1 + RQ2) established the consent-gate instrument and reported its linkability and anonymity-set readings. Two threads there were raised but not resolved, and this companion is scoped to exactly those.

(a) The unique-bridge / mix mechanism. The lead RQ2-P1 result — federation shrinks the anonymity set (ΔH < 0) — was reported honestly, but the lead paper also flagged its RQ2-P3 mechanism test as degenerate as-instrumented: the bridge-federated arm assigns a fresh bridge per circuit seed, so top-3 bridge concentration is a constant c_i = 1/C with zero variance and Spearman ρ is undefined. The mechanistic reading (developed in docs/note-unique-bridge-artifact.md) is that the adversary's observable is an exit_signature = (exit_house, bridge_label); a unique bridge per circuit makes every signature unique, so the observation-consistent anonymity set is size 1 and H≈0 by injective construction, not by funnelling. If that is right, a finite shared bridge pool should make circuits share signatures, enlarge the anonymity set, and act as a mix (concentration raises H) — the opposite of the naive funnel intuition. This makes RQ2-P3 a test of whether the lead "shrink" headline is a unique-bridge artifact that a mechanism study can qualify or correct. This mix reading connects the consent-gate bridge to the classical mix [Chaum1981] and to information-theoretic set metrics [Serjantov2002; Diaz2002] the lead paper already adopts.

(b) Churn-resilient agent selection. The lead paper held the selector at static; RQ3 asks whether an adaptive selector improves resilience when the relay pool churns. Two costs bound any such gain and are the confirmatory tension: (i) rebuilding a circuit after a dropped hop adds latency and can erode throughput, and (ii) the timing pattern of rebuilds is itself a side-channel — a rebuild-event classifier could fingerprint the selector, echoing website- and flow-fingerprinting results on onion transports [SirinamIJW18; RahmanSMGW20] and the rebuild/timing-classifier spirit of CLASI [Barton2025], and compounding statistical-disclosure exposure over repeated circuits [Danezis2003]. RQ3 therefore pairs a performance gate with an anonymity (non-fingerprint) gate: an agent selector only "helps" if it retains throughput at tolerable added latency and its rebuild pattern is not classifiable.

The lead paper's Related Work (onion routing / SOR [Egners2012], flow correlation [NasrBH18; OhYMH22], anonymity metrics [Serjantov2002; Diaz2002], social-trust G4 neighbours) is inherited unchanged. The companion adds two narrow deltas, citing only already-grounded references:

  • Bridge-as-mix vs. bridge-as-funnel. Concentrating flows through few willing bridges can be read either as a funnel (fewer distinct observation classes → smaller sets) or as a mix [Chaum1981] (shared observation class → larger sets). The set-size effect is quantified with the same entropy metrics the lead paper uses [Serjantov2002; Diaz2002]; the companion's contribution is a manipulated-concentration dose-response that adjudicates the sign, not a new estimator.
  • Rebuild-timing as a fingerprint. Churn-driven circuit rebuilds create a timing series an adversary may classify; this is the fingerprinting/timing lineage [SirinamIJW18; RahmanSMGW20; Barton2025] applied to selector-induced rebuild events rather than page loads. The companion adopts a frozen, fixture-calibrated rebuild classifier and reads its AUC as an instrument reading, mirroring the lead paper's frozen-correlator discipline (no correlator/classifier state-of-the-art is claimed).

3. Methods A — RQ2-P3 funnelling-mechanism study [FREEZE-PENDING (prereg §10 not yet signed)]

This section describes a study whose prereg (docs/rq2p3-mechanism-prereg.md) is NOT frozen. Design and parameters are operator-approved and locked; the human freeze checkpoint (SHA-256 in §10 + sign-off) is outstanding. No confirmatory cell is run until then. Everything below is the pre-registered plan; the calibration numbers cited are the pre-registered §7 gate output, explicitly not confirmatory results.

Design (manipulated-IV dose-response). A new assembler topology, bridge-federated-pool, replaces the fresh-per-seed bridge with a finite willing-bridge pool of size B under a fixed Zipf willingness skew alpha: weights = zipf_weights(B, alpha) derived from the cell seed (so the willingness profile is fixed within a run), and each circuit draws idx = weighted_draw(sha256("sor-bridge-pool|{circuit_seed}"), weights)bridge#{idx:02d}, a label reused across circuits so concentration genuinely varies. Everything downstream (hop structure, houses, exit-signature grouping, MillerMadow entropy, BCa bootstrap) is identical to the frozen lead pipeline; the lead bridge-federated branch is untouched and bit-reproducible. The manipulation grid is B ∈ {2, 4, 8} × alpha ∈ {0 (uniform), 1.0, 2.0} = 9 concentration cells; run order randomized within cell from an ordering seed distinct from the data seeds.

Hypotheses (two-sided; direction not presumed).

  • H1 (within-cell association). Spearman ρ between per-circuit top-3 willing-bridge concentration c_i and per-circuit entropy H_i. Funnel iff BCa 95% CI < 0; mix iff CI

    0; inconclusive iff CI spans 0.

  • H2 (dose-response). OLS slope β of per-run mean-H on per-run mean top-3 concentration over 9 cells × 30 runs = 270 clustered points; cell-level BCa CI; funnel iff slope CI < 0, mix iff CI > 0.
  • H3 (joint, direction-agnostic). Mechanism RESOLVED iff H1 and H2 agree in sign and both exclude 0 — the sign (funnel vs mix) is the finding; unresolved if either spans 0.

Dependent variables. Per-circuit H_i (MillerMadow entropy of the uniform posterior over the observation-consistent anonymity set, inherited verbatim) and per-circuit top-3 concentration c_i (confirm_load_rq2.bridge_concentration, unchanged).

Sampling. R = 30 seeded runs/cell, C = 50 circuits/run (matched to the lead study); base seed S0 = 20260719, per-cell seed SHA256(S0 ‖ cell_id ‖ run_index); fixed stopping rule (all 9 cells × R to completion; uninformative cell → inconclusive; no optional stopping).

Analysis. Effect size + BCa 95% CI (10,000 resamples, α = 0.05) for every test; run-level cluster bootstrap (resample whole runs, not circuits) because circuits sharing a bridge have identical c_i and correlated H_i — the same pseudo-replication defect the lead paper flagged for RQ1-P1. HolmBonferroni over this study's own family {H1-pooled, H2-slope}; the lead family-of-7 is closed and not reopened here. Any per-cell ρ contrast or ΔH replication is labelled EXPLORATORY, never a re-run of the frozen RQ2-P1.

Instrument-validation gate (§7, re-worded pre-freeze — cite docs/stage-05-rq2p3-gate-clarification.md). The §7 items were re-worded before freeze because the original items 12 encoded the naive-funnel prior and were mechanically wrong under the ratified posterior (transparent deviation logged; no hypothesis changed — H1/H2/H3 stay two-sided). The re-worded gate validates the instrument, not a sign:

  1. the frozen bridge-federated branch (not the pool) still shows the lead degeneracy — unique signatures → m_i = 1H_i ≈ 0, constant c_i = 1/C (a pool draws with replacement and cannot reproduce the injective fresh-bridge degeneracy, so the regression teeth live on the untouched branch);
  2. the B = 1 boundary yields c = 1.0 (concentration tooth) and, under the ratified posterior, H at the high end (maximal mix) — the naive "low H" gloss is refuted by construction;
  3. realized mean top-3 concentration is monotone (decreasing in B, increasing in alpha);
  4. entropy calibration inherited (H = log₂N on equiprobable synthetic senders). A §7 scope note records that the gate must not pre-assert the H-vs-concentration sign — that sign is the two-sided confirmatory question; baking it in would be funnel-circular.

Pre-registered calibration finding (NOT a confirmatory result). The dry §7 pass — synthetic, offline, no confirmatory record read — already previews a mix: across the sweep Spearman ρ runs from ≈0 up to +0.838 (all cells ρ ≥ 0), the B = 1 boundary sits at high entropy (≈2.54 bits vs the fresh-bridge reference ≈0.0), and monotonicity + entropy calibration pass. This is surfaced openly as a pre-registered calibration preview, per the §7 scope note; it does not relax the two-sided pre-commitment, and the confirmatory sign remains withheld until after freeze. Honest disclosure the eventual write-up must carry: because the dry pass already previews the mix direction, the confirmatory battery quantifies a dose-response already visible at calibration; the two-sided pre-commitment stands and the lead RQ2-P1 headline is not re-litigated.

4. Methods B — RQ3 churn-resilient agent selector (frozen prereg)

Hypotheses, gates, DVs, and analysis are frozen in sor-consent-prereg.md (§3/§4/§6) and are restated, not redefined. The two open [APPROVAL] execution params were pinned blind to RQ3 outcomes (docs/rq3-companion-run-brief.md §2).

Design. Selector strategy {static, random, agent} at the RQ3 control cell (single-house / bridge-off) under a pinned churn schedule; static is the interleaved control, and control runs are bracketed before and after the {random, agent} treatments to catch grid drift. The cells are enumerated separately from the frozen 6-cell lead lattice so the lead battery stays bit-reproducible.

Pinned execution parameters (blind).

  • Agent = qwen2.5:3b via local Ollama (agent_selector.OllamaAgentPolicy, temperature 0, per-run seed, (seed, state-hash) decision cache, deterministic heuristic fallback on query failure). Local / open-weight, $0; the Claude/frontier arm (ClaudeExploratoryPolicy) stays inert / EXPLORATORY / budget-gated and is not wired.
  • Reproducibility caveat (accepted; must be stated in the paper). Ollama at temperature 0 is not bit-identical across machines (quantization / GPU logit drift). The agent arm is reproducible via the committed decision-log + (seed, state-hash) cache replay, not via independent model re-execution on other hardware — the same honesty class as the RQ1 timing caveat. The decision log + cache are committed as the reproducibility anchor.
  • Churn = kill_prob_pct = 30, steps = 20, one deterministic schedule per run seeded from the same SHA256(S0 ‖ cell ‖ run) family; low-churn calibration baseline kill_prob_pct = 5.

Dependent variables (frozen). Throughput retention (throughput under churn / no-churn baseline); added latency = median end-to-end latency(agent) median latency(best baseline arm), in ms — a live measurement only; and rebuild-classifier AUC over the rebuild-event time series (the per-run mean inter-rebuild-gap signal), per the [Barton2025] CLASI spirit.

Confirmatory gates (frozen, family-of-7).

Test Frozen gate
RQ3-P1-perf throughput-retention(agent) max(static, random): 95% CI lower bound ≥ 10 pp
RQ3-P1-latency added-latency(agent): 95% CI upper bound ≤ 100 ms
RQ3-P2 rebuild-classifier AUC: 95% CI upper bound ≤ 0.60
RQ3-P3 logical AND: CONFIRM iff P1 ∧ P2 (perf gain without a rebuild fingerprint); else H0

R = 30 runs/cell, C = 50 circuits/run, fixed stopping rule (inherited unchanged).

Analysis + multiplicity (Holm-7 supersedes note). The three RQ3 tests were always in the frozen size-7 family {RQ1-P1, RQ1-P2, RQ2-P1, RQ2-P3, RQ3-P1-perf, RQ3-P1-latency, RQ3-P2}. Once all seven p-values exist, the companion computes the exact Holm-7 step-down over the whole family; this is the authoritative final correction and supersedes the lead paper's deliberately conservative partial embedding (7/6/5/4 report-4) — both remain valid, the partial never under-corrects, and the lead paper's already-published RQ1-P1 / RQ2-P1 survive regardless (their reported raw p ≈ 0 — a lead-paper result, not a companion figure). Effect size + BCa 95% CI for every test; nulls reported honestly (a selector that does not beat baselines, or a rebuild pattern that is classifiable, is the finding). The QUIC / ssh3 transport arm stays EXPLORATORY and deferred (design decision D3), never in the Holm family.

Calibration gates (already green; NOT confirmatory results). Two boolean gates block the RQ3 confirmatory battery and have both passed on a dry, synthetic, offline pass:

  • Churn-bites — at the pinned kp = 30 / steps = 20 the churn genuinely bites (non-zero drops/rebuilds across every RQ3 cell), so the retention and classifier tests are not degenerate.
  • Rebuild-classifier calibration — churned (kp = 30) vs low-churn baseline (kp = 5) is separable on the per-run mean inter-rebuild-gap signal (calibration AUC ≈ 0.93), while baseline-vs-baseline is not (null AUC ≈ 0.52); plus an agent cache-replay reproducibility check and the inherited entropy calibration. These are calibration readings on labelled control signals, not fit to confirmatory cells; the frozen instrument (rebuild_interval_gaps, rebuild_classifier_auc) is unchanged.

5. Results (HELD-BLIND — no confirmatory number written)

  • RQ2-P3 (H1 / H2 / H3). HELD — pending prereg freeze (§10) and the confirmatory run. No Spearman ρ, OLS slope, CI, or Holm decision is computed or inspected until after freeze. (The only figures on record are the §4/§3-A pre-registered calibration preview — ρ 0→+0.838, B=1 high-H — explicitly labelled calibration, not a result.)
  • RQ3-P1-perf / RQ3-P1-latency / RQ3-P2 / RQ3-P3. HELD — pending operator-GO on the live isolated grid. Throughput-retention margin, added-latency (live-only), and rebuild-classifier AUC over confirmatory cells are not computed; the only figures on record are the two green calibration gates (churn-bites; classifier AUC ≈0.93 sep / ≈0.52 null), labelled calibration.
  • Holm-7 (companion, authoritative). HELD until all seven confirmatory p-values exist.

6. Discussion (HELD-BLIND — placeholders)

  • Does a shared bridge funnel or mix? HELD. If H1/H2 confirm mix (ρ, slope > 0), the companion qualifies the lead "shrink" as partly a unique-bridge artifact; if funnel (< 0), the lead reading is mechanistically corroborated; if inconclusive, the mechanism stays open. The sign is the finding; the lead RQ2-P1 result is reported as-is and not re-litigated.
  • Does the agent selector help without leaking? HELD. RQ3-P3 confirms only if the agent both clears the +10 pp retention / ≤100 ms latency bar and leaves a non-classifiable rebuild pattern (AUC CI upper ≤ 0.60); either failure is an honest H0 (perf cancelled by teardown overhead, or the rebuild timing is a usable fingerprint).
  • Scope & limitations. HELD — will inherit the lead paper's lab-grid external-validity scope, add the RQ2-P3 as-instrumented concentration/mix caveat, and carry the RQ3 agent-reproducibility caveat (cache-replay, not cross-hardware re-execution) with equal prominence.

References

Inherit the lead paper's reference list (docs/stage-07-paper-draft.md §References) unchanged. The companion cites only references already grounded in the frozen sources: Chaum1981, Serjantov2002, Diaz2002, Egners2012, NasrBH18, OhYMH22, SirinamIJW18, RahmanSMGW20, Danezis2003 are in the lead paper's list; Barton2025 (CLASI rebuild/timing-classifier spirit) is grounded in the frozen lead prereg's RQ3 dependent-variable definition (sor-consent-prereg.md §3) and carries into the companion's assembled list. No reference is added that cannot be grounded from the frozen prereg or stage-01 literature.