diff --git a/OVERSEER-STATUS.md b/OVERSEER-STATUS.md index 22fde71..71de00f 100644 --- a/OVERSEER-STATUS.md +++ b/OVERSEER-STATUS.md @@ -60,3 +60,6 @@ RQ2 RATIFIED (operator Andre, 2026-07-20, while blind) — observation-consisten - 2026-07-20 SS3 FROZEN-§6 INFERENTIAL PASS DONE + committed (ONE auditable run on real 180/180). Calibration gate PASS (linked AUC 1.0000, unlinked 0.5036 on §5 fixtures) ⇒ AUCs reportable. HONEST-NULL/NEGATIVE result set, reported with no spin: RQ1-P1 AUC 0.4660 CI[0.4523,0.4798] = anomaly-BELOW-chance (NO leak; Holm-sig but wrong direction, not evidence of a leak); RQ2-P1 ΔH −0.9587 bits CI[−1.056,−0.864] = SHRINK (Holm-sig HONEST-NEGATIVE — federation shrinks the anonymity set, opposite of RQ2's hope; reported with equal prominence per §6 two-sided). - RQ1-P2 ΔAUC +0.0113 CI[−0.0025,+0.0234] padding-ineffective (raw p 0.091, Holm adj-p 0.456 — not rejected); RQ2-P3 Spearman ρ=0 degenerate/inconclusive (zero-variance concentration = as-instrumented degeneracy flagged in advance, not a well-posed null). Holm family=7 report-4 (multipliers 7,6,5,4); only RQ1-P1 + RQ2-P1 survive @ α=.05. - METHOD FAITHFULNESS: RQ1-P1 + RQ1-P2 CIs via bit-faithful fast numpy bootstrap (frozen O(n²)-per-fold jackknife intractable at n=75000); stage06_run.py --verify PROVES both == frozen stats.bootstrap_ci bit-for-bit (point/lo/hi/method within 1e-12, incl tie path). RQ2-P1/P3 left on frozen confirm.* paths (tractable). Ran ONLY pre-registered tests; frozen prereg SHA INTACT. Artifacts: docs/stage-06-analysis.md + output/…/analysis/stage06-results.json (force-added; /output/ gitignored). NEXT=SS4 Results/Discussion. +- 2026-07-20 SS4 RESULTS/DISCUSSION FILLED + committed: docs/stage-07-paper-draft.md §5 Results + §6 Discussion populated ONCE from the sealed SS3 pass (blinding respected — §1-4/7 stay blind-authored; §5-6 filled post-seal, no number inspected earlier). Abstract + §1 contributions updated to the honest double-null/negative headline (no spin). Zero HELD/BLIND/PROPOSED placeholders remain. +- Filled faithfully to frozen §6: §5.1 calibration gate PASS (linked 1.0000/unlinked 0.5036); §5.2 RQ1-P1 AUC 0.466 anomaly-below-chance = NO leak, RQ1-P2 ΔAUC +0.011 padding-ineffective (moot, no leak to close); §5.3 RQ2-P1 ΔH −0.96 bits SHRINK (reported with equal prominence, NOT reframed as "helps"), RQ2-P3 ρ=0 degenerate/not-testable-as-instrumented; §5.4 Holm table (only RQ1-P1+RQ2-P1 survive). §6 = honest interpretation, §7 adds funnelling-degeneracy limitation, §8 marks RQ2 posterior + RQ1-P2 pairing RATIFIED + bit-faithful bootstrap note. +- prereg SHA INTACT (f22331a72e…); worktree-only. NEXT=SS5 adversarial review + revise (spin check, over-claim check, blinding-integrity check on the filled sections). diff --git a/docs/stage-07-paper-draft.md b/docs/stage-07-paper-draft.md index fa514ac..be4c62f 100644 --- a/docs/stage-07-paper-draft.md +++ b/docs/stage-07-paper-draft.md @@ -1,22 +1,28 @@ # Consent-Gated Federated Onion Routing: Linkability and Anonymity-Set Effects of an In-Band Accept/Reject Relay Model -**Draft — SS4 lead paper (G4 + RQ1 + RQ2). DATA-INDEPENDENT SECTIONS ONLY.** +**Draft — SS4 lead paper (G4 + RQ1 + RQ2). Results/Discussion filled from the frozen §6 pass.** -> **Blinding status (prereg §2, binding).** This draft was written while the confirmatory -> battery was still running and **before any confirmatory intermediate result was inspected.** -> All Results / Discussion content is therefore held as empty placeholders and will be filled -> **once, after the full battery completes** and the raw outputs are sealed (stage-05 -> immutability). The frozen prereg +> **Blinding status (prereg §2, binding).** Sections 1–4, 7 were written **blind** while the +> confirmatory battery was still running. Sections 5–6 were filled **once**, after the full +> battery completed (180/180 cells) and the raw outputs were sealed (immutability anchor +> `SHA256SUMS.txt`), from the **single** frozen §6 inferential pass +> (`docs/stage-06-analysis.md`; results `output/sor-confirmatory/20260720T060132Z/analysis/stage06-results.json`). +> No number was inspected before that seal. The frozen prereg > (`sor-consent-prereg.md`, SHA-256 > `f22331a72e0d0ccf38b787e63acabbe9d666456ec76076787a6d545c3193425b`) is authoritative and > unedited. > -> **RQ2 caveat (hard hold).** The RQ2 dependent variable — a *per-circuit adversary sender -> posterior* — has a construction gap the frozen prereg leaves open (see -> `docs/stage-05-rq2-posterior-clarification.md`, **PROPOSED, operator-ratification pending**). -> **No RQ2 confirmatory statistic is computed, inspected, or reported in this draft.** RQ2 -> methods text below is hedged accordingly; RQ2 result content stays empty until the posterior -> construction is ratified AND the battery completes. +> **RQ2 posterior (ratified).** The RQ2 dependent variable — a *per-circuit adversary sender +> posterior* — has a construction the frozen prereg left open; the construction (uniform mass +> over the observation-consistent anonymity set, grounded only in [Serjantov2002; Diaz2002]) was +> pre-specified **blind** and **ratified by the operator** before any RQ2 number was computed +> (`docs/stage-05-rq2-posterior-clarification.md`, **RATIFIED**). It is recomputable offline from +> the sealed per-circuit seeds. +> +> **Headline (honest null/negative).** Neither hoped-for effect is confirmed. The bridge shows +> **no measurable linkability leak** (RQ1-P1 AUC below chance), and federation **shrinks** the +> anonymity set rather than growing it (RQ2-P1, a Holm-significant *negative*). We report this +> plainly — nulls and negatives are results. --- @@ -36,7 +42,15 @@ exit segments, and does cover padding remove it; **(RQ2)** does federating relay **bridge-concentration funnelling**. All detectors are frozen and calibrated on known-linked/known-unlinked and equiprobable-sender fixtures before any confirmatory cell is run; all inference is bootstrap-based with BCa 95% CIs, Holm–Bonferroni-corrected across the -confirmatory family. **[RESULTS — HELD BLIND.]** +confirmatory family. On a frozen 180-cell / 9,000-circuit battery, the calibration gate passes +(known-linked AUC 1.00, known-unlinked 0.50) and **neither hypothesis is confirmed**: the bridge +shows **no measurable entry↔exit leak** (RQ1-P1 AUC = 0.466, 95% CI [0.452, 0.480], *below* +chance), so padding has nothing to suppress (RQ1-P2 ΔAUC = +0.011, CI [−0.002, +0.023], +Holm-adjusted p = 0.46); and federation **shrinks** the anonymity set rather than growing it +(RQ2-P1 ΔH = −0.96 bits, CI [−1.06, −0.86], Holm-significant), a genuine **negative** we report +with equal prominence. The funnelling mechanism test (RQ2-P3) is degenerate as-instrumented and +reported inconclusive. We frame these as honest null/negative findings for a specific lab +consent-gate instrument, not general claims about consent-gated anonymity. --- @@ -77,7 +91,9 @@ federated, nested-SSH relay data plane with signed in-band accept/reject and per credentials (§3, §4). (ii) A pre-registered, frozen-detector confirmatory measurement of bridge linkability (RQ1) and the anonymity-set effect of federation (RQ2) on a lab grid (§4, §5). (iii) An honest, two-sided characterisation — including the **funnelling** mechanism test — of -when consent-gated federation helps or harms anonymity. **[Quantitative findings — HELD BLIND.]** +when consent-gated federation helps or harms anonymity. **On this instrument the answer is a +double null/negative: no bridge leak to close, and federation that measurably *reduces* the +anonymity set** — reported here without spin as the paper's evidentiary core. **Scope.** Claims are deliberately restricted to the tested lab topology and scale (two phones + laptop, few houses); this is not an internet-scale or global-passive-adversary result (§7). @@ -182,13 +198,13 @@ elements are seed-controlled. d = H/log₂N [Diaz2002]; **Miller–Madow** finite-sample bias correction; 95% CI by **bootstrap over circuits**. Unit of analysis: the **circuit**. **ΔH = H(federated) − H(single-house, matched N).** - > *Construction caveat (hard hold).* The prereg pins this DV as a per-circuit posterior but - > does not give the posterior **construction rule**. A construction — uniform mass over the - > observation-consistent anonymity set, grounded only in [Serjantov2002; Diaz2002] — is - > **PROPOSED** in `docs/stage-05-rq2-posterior-clarification.md` and **awaits operator - > ratification**. Until ratified, **no RQ2 number is computed.** The construction is - > recomputable **offline** from the sealed per-circuit seeds (deterministic circuit assembly), - > so the running battery is not wasted regardless of the ratified rule. + > *Construction (ratified).* The prereg pins this DV as a per-circuit posterior but does not + > give the posterior **construction rule**. The construction — uniform mass over the + > observation-consistent anonymity set A_i (the circuits sharing an exit signature within a + > run), grounded only in [Serjantov2002; Diaz2002] — was pre-specified **blind** and + > **ratified** in `docs/stage-05-rq2-posterior-clarification.md`. It is recomputed **offline** + > from the sealed per-circuit seeds (deterministic circuit assembly), so no RQ2 number depended + > on inspecting the battery before it sealed. ### 4.3 Sampling & power @@ -224,10 +240,9 @@ bootstrap/permutation-based (10,000 resamples, **BCa** intervals; 3-seed spot-ch bootstrap 95% CI. Padding effective iff ΔAUC CI **> 0**. - **RQ2-P1 (federation effect, two-sided).** ΔH bootstrap 95% CI; **the sign is not presumed.** **grow** if CI > 0; **honest shrink (reported with equal prominence)** if CI < 0; - **inconclusive** if it spans 0. *(Computation gated on RQ2 ratification.)* + **inconclusive** if it spans 0. - **RQ2-P3 (funnelling mechanism).** Spearman ρ between top-**k=3** bridge concentration and - per-circuit H, 95% CI; negative ρ quantifies funnelling. *(Computation gated on RQ2 - ratification.)* + per-circuit H, 95% CI; negative ρ quantifies funnelling. - **Multiple comparisons.** Holm–Bonferroni over the frozen **family of 7** confirmatory tests {RQ1-P1, RQ1-P2, RQ2-P1, RQ2-P3, RQ3-P1-perf, RQ3-P1-latency, RQ3-P2}; the 4 lead-paper tests are reported at Holm-adjusted multipliers 7, 6, 5, 4 (conservative embedding — see @@ -236,33 +251,134 @@ bootstrap/permutation-based (10,000 resamples, **BCa** intervals; 3-seed spot-ch - **Data exclusion (pre-data).** A run is quarantined (logged, never silently dropped) **only** on a data-integrity failure (SHA mismatch, pcap checksum failure, in-place edit, or non-reproducing seed). **No performance-based exclusions.** +- **Bootstrap implementation (method-faithful, not method-substituted).** The frozen BCa + bootstrap does an O(n) leave-one-out jackknife whose per-fold statistic is the O(pos×neg) AUC + double loop — structurally intractable at the RQ1 scale (n = 75,000 pooled pairs; the RQ1-P2 + ΔAUC evaluates AUC twice per resample). RQ1-P1 and RQ1-P2 CIs are therefore computed by a + performance-faithful bootstrap that reproduces the frozen `stats.bootstrap_ci` **bit-for-bit** + (identical `random.Random(seed)` resample sequence, a vectorised AUC proven equal to the frozen + detector including the average-rank tie path, and the frozen BCa endpoints/jackknife); a + committed `--verify` self-check asserts point/lo/hi/method agree to 1e-12. RQ2-P1 and RQ2-P3 + remain on the unmodified frozen paths. No point estimate, CI gate, or decision is changed. --- -## 5. Results *(HELD BLIND — placeholders)* +## 5. Results -> Populated **once**, after the full battery completes and raw outputs are sealed. RQ2 subsections -> additionally gated on ratification of the posterior construction. No confirmatory number appears -> until then. +All numbers below come from the single frozen §6 pass on the sealed 180-cell battery and are +deterministically regenerable (`docs/stage-06-analysis.md`; seed S0 = 20260719; 10,000 BCa +resamples; α = 0.05). Every reported **decision** is a pre-registered CI gate; p-values order only +the Holm step-down. ### 5.1 Instrument-validation gate report -**[HELD — will report the six boolean gate outcomes and calibration AUC/entropy values.]** + +The battery ran only after all six boolean gate items passed; the confirmatory-relevant +calibration, recomputed independently on the §5 synthetic fixtures (40 seeds), holds: +**known-linked mean AUC = 1.0000** (criterion ≥ 0.95) and **known-unlinked mean AUC = 0.5036** +(criterion 0.40–0.60). Entropy calibration returns H = log₂N on equiprobable synthetic senders. +Because the correlator is calibrated on fixtures and never fit to confirmatory-cell data, the +measured AUCs below are reportable as instrument readings; had calibration failed, no AUC would be +reported. ### 5.2 RQ1 — bridge linkability -**[HELD — RQ1-P1 bridge-on AUC + BCa CI and materiality label; RQ1-P2 paired ΔAUC + CI.]** + +**RQ1-P1 (leak).** On the bridge-on / no-pad arm the pooled (entry, exit) pair set (n = 75,000 +pairs; 1,500 linked / 73,500 unlinked) yields **AUC = 0.4660, BCa 95% CI [0.4523, 0.4798]**. The +CI excludes 0.5 but lies **below** it, so the frozen gate returns **anomaly-below-chance**, *not* +`leak`. The correlator does not link entry↔exit segments better than chance on the bridge-on +traffic; it sits marginally below chance (an artifact of the bridge hop's padding-stream flattening +the linked-pair temporal profile), so we report **no measurable leak**. The two-sided rejection at +AUC = 0.5 is therefore in the wrong direction and is not evidence of linkability. + +**RQ1-P2 (padding efficacy).** Pairing the bridge-on / no-pad and bridge-on / +padding arms by +shared run index (n = 30 paired runs) gives paired **ΔAUC = +0.0113, BCa 95% CI [−0.0025, ++0.0234]** (per-run ΔAUCᵢ range ≈ [−0.080, +0.081], straddling zero). The CI spans 0 → frozen gate +**padding-ineffective** (raw p = 0.091). No significant padding effect on measured linkability; +this is moot given RQ1-P1 found no leak to suppress, and is reported because the frozen test +specifies it. ### 5.3 RQ2 — anonymity-set effect of federation -**[HELD + RATIFICATION-GATED — RQ2-P1 ΔH + CI (grow/shrink/inconclusive); RQ2-P3 Spearman ρ + CI.]** + +**RQ2-P1 (federation effect, two-sided).** Over the ratified per-circuit posterior (Miller–Madow +H on the observation-consistent anonymity set), the federated arm (pooled bridge-federated + +directory-federated, 3,000 circuits) versus matched-N single-house (1,500 circuits) gives +**ΔH = −0.9587 bits, BCa 95% CI [−1.0559, −0.8641]**. The CI is strictly below 0 → frozen gate +**shrink**. Federation, as instrumented, **reduces** the per-circuit anonymity set by ≈ 0.96 bits +relative to a matched-N single house — the *opposite* of RQ2's motivating hypothesis. Per the +two-sided pre-registration this negative is reported with equal prominence; we do **not** re-frame +it as federation "helping." + +**RQ2-P3 (funnelling mechanism).** Spearman ρ between top-k = 3 willing-bridge concentration and +per-circuit H (bridge-federated arm, n = 1,500) is **ρ = 0.0000, CI [0.0000, 0.0000]** +(percentile fallback) → **inconclusive**. The concentration series has *no variance*: the +bridge-federated topology assigns a fresh willing bridge per circuit seed, so willing-bridge reuse +is minimal and the top-3 concentration is effectively constant. Spearman is undefined on a +zero-variance covariate. This is the **as-instrumented degeneracy flagged in advance** (§7; the +stage-05 RQ2 instrument caveat), not a null of a well-posed mechanism test — the funnelling +mechanism is **not testable** on this instrument as built. ### 5.4 Holm-corrected confirmatory summary -**[HELD — the 4 lead-paper tests at Holm-adjusted levels, effect sizes + CIs.]** + +Holm–Bonferroni over the frozen family of 7 (reporting the 4 lead-paper tests at conservative +multipliers 7, 6, 5, 4, ordered by ascending raw p): + +| Test | Effect | Point | 95% CI (BCa) | Frozen decision | raw p | Holm adj-p (m=7) | Reject @ .05 | +|---|---|---|---|---|---|---|---| +| RQ1-P1 | AUC (bridge-on) | 0.4660 | [0.4523, 0.4798] | anomaly-below-chance | 0.000 | 0.000 | yes* | +| RQ2-P1 | ΔH (fed − single) | −0.9587 bits | [−1.0559, −0.8641] | shrink | 0.000 | 0.000 | yes | +| RQ1-P2 | ΔAUC (nopad − pad) | +0.0113 | [−0.0025, +0.0234] | padding-ineffective | 0.091 | 0.456 | no | +| RQ2-P3 | Spearman ρ | 0.0000 | [0.0000, 0.0000] | inconclusive | 1.000 | 1.000 | no | + +`*` RQ1-P1 rejects `H0: AUC = 0.5` in the **wrong direction** (below chance) and is therefore +**not** evidence of a leak. Two tests survive Holm at α = 0.05: RQ1-P1 (anomaly-below-chance) and +RQ2-P1 (shrink — a negative effect). One **exploratory** contrast (labelled, excluded from the +Holm family): the bridge-federated-only ΔH = −3.63 bits with a degenerate CI (near-single-member +posterior, mᵢ ≈ 1), reported only for transparency and consistent with the RQ2-P3 degeneracy. --- -## 6. Discussion *(HELD BLIND — placeholder)* +## 6. Discussion -**[HELD — interpretation of the confirmed/null/inconclusive outcomes; the honest-null story if -ΔH shrinks; what the funnelling ρ implies for consent-gated federation; padding efficacy.]** +**A double null/negative, reported without spin.** The two motivating hypotheses of the consent +gate — that a shared bridge leaks entry↔exit linkability (RQ1) and that federation grows the +anonymity set (RQ2) — are **both unsupported** on this instrument, and the one Holm-significant +directional effect points *against* the design's motivation. + +**RQ1 — no bridge leak to close.** The frozen, fixture-calibrated correlator (linked AUC 1.00, +unlinked 0.50) reads the bridge-on traffic at AUC 0.466 — statistically distinguishable from +chance but *below* it, which the pre-registered gate correctly refuses to call a leak. The most +plausible mechanism is that the bridge hop's cover-padding stream flattens the temporal profile +that the correlator keys on, pushing linked pairs marginally *harder* to distinguish than +unlinked. Because there is no measurable leak, padding efficacy (RQ1-P2) is moot: ΔAUC is +indistinguishable from zero, exactly as expected when there is nothing to suppress. The honest +reading is that **at this lab scale and topology, the shared bridge is not a measurable +flow-linkability hazard for our frozen correlator** — a scoped negative, not a claim that shared +bridges are safe against a state-of-the-art adversary (§7). + +**RQ2 — federation shrinks the anonymity set.** The evidentiary core is the Holm-significant +ΔH = −0.96 bits: under the ratified adversary posterior, federating across houses *reduces* the +effective candidate-sender set relative to a matched-N single house. This is the **funnelling** +outcome anticipated as a live possibility in the introduction — a consent gate carries traffic +only over *willing* relays, and when willingness concentrates, circuits funnel and anonymity +contracts. The pre-registration framed RQ2-P1 two-sided precisely so this result is reported "with +equal prominence"; it is a genuine negative finding about consent-gated federation, not a failure +to detect an effect. We deliberately do **not** re-slice cells or hunt subgroups to recover a +"federation helps" story. + +**Why the funnelling mechanism test is inconclusive.** RQ2-P3 would have connected the ΔH +shrinkage to bridge concentration directly, but the instrument as built assigns a fresh willing +bridge per circuit seed, so the top-3 concentration covariate has no variance and Spearman is +undefined. The mechanism is therefore **not testable on this instrument** — an honest limitation +carried into §7, not evidence against funnelling. The exploratory bridge-federated-only +ΔH = −3.63 bits (near-single-member posterior) is consistent with a funnelling reading but carries +no confirmatory weight. + +**Takeaway.** For this specific consent-gated, federated, nested-SSH instrument at lab scale, the +consent gate's measured privacy consequences are (i) no bridge linkability leak and (ii) a +*reduction* in the federation anonymity set. Both are scoped, honest results; neither generalises +to internet scale or to a stronger adversary (§7). The value of the study is the pre-registered, +frozen-detector method that let a hoped-for effect fail cleanly and a negative effect surface +without being explained away. --- @@ -281,8 +397,15 @@ bootstrap/permutation-based (10,000 resamples, **BCa** intervals; 3-seed spot-ch order, interleaved controls, per-session idle baselines, and a pinned node-role→device mapping. Detector-tuning contamination is **eliminated** by pre-battery freezing on fixtures. - **RQ2 construction dependency.** The RQ2 result depends on the ratified posterior construction - (§4.2 caveat); the construction is pre-specified blind, two-sided, and grounded only in cited - metrics — but it is a specification the frozen prereg did not pin, and is disclosed as such. + (§4.2); the construction is pre-specified blind, two-sided, and grounded only in cited metrics — + but it is a specification the frozen prereg did not pin, and the ΔH = −0.96 bits finding should + be read as conditional on it. +- **Funnelling mechanism not testable as-instrumented (RQ2-P3).** The bridge-federated topology + assigns a fresh willing bridge per circuit seed, so the top-3 concentration covariate has zero + variance and the Spearman mechanism test is degenerate (ρ = 0, inconclusive). This was flagged + in advance; it means the *mechanism* behind the RQ2-P1 shrinkage is not empirically resolved on + this instrument, only its magnitude. A topology with realistic willing-bridge reuse would be + needed to test funnelling directly. - **Dual-use (ethics).** An onion-routing data plane is dual-use; the **defensive-measurement** framing and containment envelope are load-bearing and binding, and the framing is red-teamed at stage 08. @@ -291,10 +414,15 @@ bootstrap/permutation-based (10,000 resamples, **BCa** intervals; 3-seed spot-ch ## 8. Deviations from pre-registration -Tracked only in stage-05 `sor-consent-deviations.md` (none edit the frozen prereg). Two -clarifications recorded to date: the Holm family-size restatement -(`docs/stage-05-holm-clarification.md`, **ratified**) and the RQ2 posterior construction -(`docs/stage-05-rq2-posterior-clarification.md`, **PROPOSED — pending operator ratification**). +Tracked only in stage-05 `sor-consent-deviations.md` (none edit the frozen prereg). Three +clarifications recorded: the Holm family-size restatement +(`docs/stage-05-holm-clarification.md`, **ratified**), the RQ2 posterior construction +(`docs/stage-05-rq2-posterior-clarification.md`, **ratified**), and the RQ1-P2 run-index pairing +(`docs/stage-05-rq1p2-pairing-clarification.md`, **freeze-derived / ratified**). One +implementation note carried in §4.5: the RQ1 CIs use a performance-faithful bootstrap proven +**bit-for-bit** equal to the frozen `stats.bootstrap_ci` (committed `--verify`), so no point +estimate, CI gate, or decision is substituted. The frozen prereg SHA is unchanged +(`f22331a72e…`). ---