Make a clean clone reproducible without a sci-method checkout: - vendor the frozen pre-registration (docs/prereg/, pinned SHA-256 f22331a7… unchanged) and the 36-source bibliography into the repo - confirmatory_run.py resolves the in-repo prereg first, falling back to the canonical sci-method path; the SHA gate is identical (the pin, not the path, is the guarantee) - add REPRODUCE.md: tiered clone→env→seed→run→analyze checklist, honest about the isolated-engine containment rule and the 1.2 GB raw data that is not committed Defensive-measurement instrument; no change to measured behavior. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
8.3 KiB
Reproducing the sor-consent study
A step-by-step checklist for reproducing the linkability / anonymity-set / churn-resilience
measurements reported in docs/stage-07-paper-draft.md and docs/stage-07-companion-methods.md.
This instrument is built for defensive measurement of a trust model's exposure. It moves only
self-generated fixture traffic on our own lab grid, and every forwarder runs inside an
isolated engine (docker / multipass) — never on the host (assert engine != local or the code
refuses). Reproduce it in the same spirit: no external targets, no third-party traffic.
Self-contained. Everything needed to reproduce is committed in this repo — no sci-method
checkout is required. The reproduction deposits are:
| Deposit | Path | Purpose |
|---|---|---|
| Frozen pre-registration | docs/prereg/sor-consent-prereg.md (+ .sha256) |
the hashed, frozen design; the launcher SHA-checks it |
| Bibliography | docs/sor-consent-bibliography.md |
full 36-source reference list for the paper |
| Analysis + instrument | cmd_chat/sor/ |
every DV measure and analysis pass lives here |
| Analysis results + integrity manifest | output/sor-confirmatory/<ts>/ — analysis/stage06-results.json, SHA256SUMS.txt, INTEGRITY-REPORT.md, provenance.json |
the committed, verifiable outputs of the frozen run |
Raw data is not committed. The per-run pcaps + event logs (
output/sor-confirmatory/<ts>/confirmatory-data/, ~1.2 GB: 9,000events.jsonl+ 27,000 pcaps) are.gitignored for size. What is committed is theirSHA256SUMS.txtmanifest and the analysis results. To re-run the analysis from raw pcaps you must either obtain that archive out-of-band or re-collect it via §5; a clean clone can still verify the committed results and the integrity manifest.
0. Where does the code run? (the common confusion)
| Component | Runs where | How |
|---|---|---|
| Clone, venv, orchestration, analysis | On the host (your laptop) | you drive everything from the repo checkout |
| Circuit forwarders / hops | Inside an isolated engine (docker container or multipass VM) | spawned automatically by the code via the isolation prefix (cmd_chat/agent/bridge.py:517-530); the host never forwards |
You clone once, on the host. You do not clone inside the VM. When a run starts, each hop is
launched into its own container/VM for you. There is a hard containment check: if the engine
resolves to local, the forwarder refuses to run.
1. Get the code
git clone https://git.churchofmalware.org/trilltechnician/hack-house.git
cd hack-house
git checkout feat/sor-consent-relay # the branch this study lives on
2. Set up the environment (host)
hh/scripts/bootstrap.sh # Python venv + server deps + Rust client build
# (or: python3 -m venv .venv && . .venv/bin/activate && pip install -r requirements.txt)
Requires Python 3.10+ and a Rust toolchain (bootstrap builds the client).
Install at least one isolated engine — forwarders will not run without one:
hh/scripts/ensure-docker.sh # docker (the default engine, --engine docker)
hh/scripts/ensure-multipass.sh # multipass (full-VM isolation), optional
3. Verify the instrument (no grid, no confirmatory data)
pytest tests/ # Python acceptance checks (R1–R7)
cd hh && cargo test && cd .. # Rust `_sor` consent-parser proptests
# Safe start-line preflight — runs the §5 gate, grid inventory, cell plan,
# a 2-run deterministic dry provenance pass, and the assembler dry check.
# Collects NO confirmatory data; prints a GO/NO-GO summary.
python -m cmd_chat.sor.confirmatory_run
A green preflight means the six instrument-validation gate items pass: 3-hop end-to-end delivery
with per-hop pcap + checksum, seeded reproducibility, correlator calibration (linked AUC≈1 /
unlinked≈0.5), entropy calibration (H=log₂N), isolated-engine-only forwarders, and append-only
provenance whose events.jsonl SHA-256 matches the manifest.
4. Reproduce the reported analysis
Every number in the paper is deterministically regenerable from the frozen raw data and the single base seed S0 = 20260719 — same seed → identical output.
Two things a clean clone can always do (no grid, no raw data): re-hash the committed results and verify the raw-data integrity manifest:
cd output/sor-confirmatory/20260720T060132Z && sha256sum -c SHA256SUMS.txt # needs raw data present
# and inspect the committed frozen outputs directly:
cat output/sor-confirmatory/20260720T060132Z/analysis/stage06-results.json
To regenerate the numbers from raw pcaps you need the ~1.2 GB confirmatory-data/ archive
present (obtain out-of-band, or re-collect via §5). With it in place:
# Lead paper — RQ1-P1, RQ1-P2, RQ2-P1, RQ2-P3 (Holm family_size=7, report 4)
python -m cmd_chat.sor.analysis.stage06_run \
output/sor-confirmatory/20260720T060132Z/confirmatory-data \
--out output/sor-confirmatory/20260720T060132Z/analysis/stage06-results.json
# Companion — RQ2-P3′ mechanism study and RQ3
python -m cmd_chat.sor.analysis.rq2p3_confirm # shared-pool anonymity-set mechanism
python -m cmd_chat.sor.analysis.rq3_confirm # churn-resilience selector
Same seed → identical output (10,000 BCa bootstrap resamples, α = 0.05). Verify integrity by
re-hashing each run's events.jsonl against its manifest.json, and confirming the frozen prereg
still hashes to f22331a72e0d0ccf38b787e63acabbe9d666456ec76076787a6d545c3193425b.
5. Re-run the full confirmatory battery (human-gated — needs the lab grid)
The immutable data run (R = 30 runs × C = 50 circuits = 9,000 live circuits) is deliberately human-gated. It is triple-locked and refuses to launch unless all four hold:
- Operator token —
SOR_CONFIRMATORY_GO=1in the environment. - Frozen prereg — the on-disk prereg hashes to
f22331a7…(see §6 caveat). - Isolation —
--engine docker(ormultipass); neverlocal. - Full grid — every device reachable (the launcher refuses a degraded grid rather than fabricate cells).
SOR_CONFIRMATORY_GO=1 python -m cmd_chat.sor.confirmatory_run --operator-go --engine docker
python -m cmd_chat.sor.rq3_confirmatory_run --engine docker # companion RQ3 battery
Then re-run the §4 analysis over the freshly collected confirmatory-data/.
Seeds & determinism
- Base seed:
S0 = 20260719(fixed pre-data). - Per-cell seed:
SHA256(S0 ‖ cell_id ‖ run_index)— drives every stochastic element (path selection, churn schedule, padding jitter) so two runs with the same seed produce an identical circuit-build sequence. - Seed + git SHA + node-role→device mapping are echoed into each run's immutable
manifest.json.
Outputs (treat as immutable once written)
output/sor-runs/<ts>/ manifest.json · events.jsonl · pcap/hop<N>.pcap · metrics.json
output/sor-confirmatory/<ts>/ preflight artifacts · confirmatory-data/ · analysis/
Honest reproduction caveats
- The frozen prereg is vendored. The confirmatory launcher (§5) hard-refuses unless the
pre-registration is present and hashes to
f22331a7…. The frozen prereg is committed in-repo atdocs/prereg/sor-consent-prereg.md(with.sha256), so this checkout is self-contained — nosci-methodaccess is needed. The launcher prefers the in-repo copy and falls back to the operator's canonicalsci-methodcheckout; the SHA-256 pin (not the path) is the integrity guarantee. Verify it yourself:cd docs/prereg && sha256sum -c sor-consent-prereg.sha256. - The physical grid is lab-specific. The reported run distributes hops across our own 2 phones
- laptop + hack-house VM churn fabric. An external reproducer instead uses the containerised-hop fallback (all hops as isolated docker containers) — this reproduces the instrument and the analysis, but not the exact physical device distribution.
- Agent arm (RQ3). Local open-weight (Ollama, temp 0) is not bit-identical across machines
(quantization / GPU logit drift). The agent arm reproduces via the committed decision-log +
(seed, state-hash)cache replay, not via independent cross-hardware model re-execution.
See CLAUDE.md for the containment law and docs/stage-07-paper-draft.md §4 for the frozen method.