df8f1881d8
Fourth benchmark axis: teams of LLM agents deliberate in a room, implement code in an isolated VM, and are scored deterministically on correctness/speed. - Multi-language adapter (python/js/go/rust/bash) via MultiPL-E continuation mode - Append-only JSONL ledger with status tracking (ok/dnf/killed/error) so budget-exhausted or crashed runs still record a row (fixes selection bias) - Model-aware wall-clock scaling (U-shaped by param count; 3x for reasoning) - Self-owned SIGALRM/SIGTERM watchdog (RunTimeout: BaseException so broad except Exception handlers in the infer/completion path can't swallow it) - Seed forwarded to Ollama sampler + markdown-fence stripping in completion.py
484 lines
27 KiB
Markdown
484 lines
27 KiB
Markdown
# Agent Olympics — a hackathon-competition benchmark inside hack-house
|
||
|
||
> Status: **DRAFT spec** (design only — no implementation yet). Author: bench team.
|
||
> Scope: a fourth benchmark axis layered on `hh/scripts/bench/`. Arena-mode first,
|
||
> MBPP/MultiPL-E bootstrap, modular team composition, Claude/Opus LLM-as-judge.
|
||
> This document is the build contract. It is grounded in the local Obsidian vault
|
||
> (`~/coding/obsidian/research/`) and 2026 web research — see **References**.
|
||
|
||
---
|
||
|
||
## 1. One-paragraph concept
|
||
|
||
Stand up **two (or more) teams of LLM agents** inside hack-house. Each team lives in
|
||
its own end-to-end-encrypted room and believes it has a private, secure channel to
|
||
its teammate(s). Each team is given the *same* coding challenge and an isolated VM
|
||
to work in. Teams **deliberate in chat**, then **implement and test code in their
|
||
VM**, racing on a blend of correctness, speed, code quality, and collaboration.
|
||
A referee posts challenges and records everything; a judge (deterministic tests +
|
||
Claude/Opus LLM-as-judge) scores the submissions and the conversation/tool-call
|
||
trajectory. Aggregate across a slate of events into a **medal table**. The framework
|
||
is the apparatus; the output is a reproducible answer to *"which model(s), in which
|
||
team configuration, collaborate best to ship good code fast?"*
|
||
|
||
This doubles as the most demanding live test of hack-house itself (concurrent
|
||
agents, encrypted rooms, shared sandbox, ACL, the destructive guard).
|
||
|
||
---
|
||
|
||
## 2. Goals & non-goals
|
||
|
||
### Goals
|
||
- **G1 — Real comms substrate.** Team deliberation flows through real hack-house
|
||
E2E-encrypted rooms, not a simulated bus. The product is the medium.
|
||
- **G2 — Modular team composition.** Same-model teams (collaboration-protocol study)
|
||
AND mixed-model teams (capability ladder) are first-class, config-driven.
|
||
- **G3 — Bootstrap fast.** Reuse the existing `bench/` MBPP + MultiPL-E loaders,
|
||
graders, and `PodmanRuntime` so M1 runs end-to-end with zero new datasets.
|
||
- **G4 — Extensible challenges.** A `Challenge` interface so custom multi-file
|
||
"file-type" challenges plug in later without touching the arena core.
|
||
- **G5 — Trustworthy scoring.** Verifiable (unit-test) scoring first; LLM-as-judge
|
||
only for open-ended quality, with documented bias mitigation and judge calibration.
|
||
- **G6 — Reproducible & auditable.** Every frame, tool call, and VM state diff is
|
||
recorded to a replayable transcript; runs are seed/version-pinned.
|
||
- **G7 — Measure the placebo.** Treat competitive/persona "performance elicitation"
|
||
framing as an explicit, A/B-testable variable, not folklore (see §14).
|
||
- **G8 — Safe by construction.** Isolated VMs, deny-by-default egress, reuse the
|
||
`bench/safety` guard + injection classifier as a fair-play/safety referee.
|
||
|
||
### Non-goals (for now)
|
||
- Not editing `cmd_chat/` (tool code). Arena mode drives turns externally and posts
|
||
into real rooms. Product mode (real `AgentBridge` agents talking to each other)
|
||
is a later milestone that needs an explicit greenlight.
|
||
- Not a training loop (no RL/fine-tuning). This selects and compares; it does not train.
|
||
- Not a public leaderboard. Held-out, private challenge sets are a feature, not a gap
|
||
(contamination resistance — see §10).
|
||
|
||
---
|
||
|
||
## 3. Prior art this design borrows from (grounding)
|
||
|
||
- **Read-vs-write rule for multi-agent work.** Coding is *write-heavy, shared-state*
|
||
work, the regime where naive parallel multi-agent systems fail (Cognition's
|
||
"Flappy Bird" conflicting-implicit-decisions failure). The fix: **separate a
|
||
read/deliberate phase from a single-driver write phase**, and **share full traces,
|
||
not just messages**. We bake this into the loop (§7) *and* measure the coordination
|
||
tax. [[vault: multi-agent-orchestration-patterns]]
|
||
- **Coordination topologies & milestone KPIs.** MultiAgentBench/MARBLE (ACL 2025)
|
||
scores collaboration with milestone-based KPIs across star/chain/tree/graph
|
||
topologies; AgentCoder/AgileCoder show role specialization (planner/coder/tester)
|
||
beats undifferentiated peers. We make role + topology config knobs. [web]
|
||
- **Eval-harness discipline.** Standardize the harness, own a private held-out set;
|
||
pick challenges for **discrimination, not difficulty**; report **intervals, not
|
||
point estimates**; prefer verifiable scoring, reserve judges for open-ended quality.
|
||
[[vault: eval-harnesses-benchmark-design]]
|
||
- **Outcome vs trajectory.** Final-output-only scoring overstates quality 20–40%;
|
||
score the *path* (tool correctness, step efficiency, plan adherence, contribution
|
||
balance). Use **agent-as-a-judge** to walk the trajectory. Report **pass^k**
|
||
(worst-case) beside pass@k for reliability. [[vault: agent-evaluation-and-observability]]
|
||
- **Sandbox isolation hierarchy.** Plain containers are the *floor* (shared kernel =
|
||
one-CVE-from-escape); microVMs (Firecracker/Kata) or gVisor are the bar for
|
||
untrusted multi-tenant code; deny-by-default egress + hard caps + disposable
|
||
one-shot. Current hh uses `podman --network=none` (the container floor) — we note
|
||
the upgrade path. [[vault: agent-sandboxing-isolation]]
|
||
- **LLM-as-judge biases.** Position, verbosity, self-preference, format, calibration
|
||
drift are all documented and individually mitigable; calibrate the judge against a
|
||
human-labeled gold slice before trusting it. [web + [[vault: llm-evals]]]
|
||
|
||
---
|
||
|
||
## 4. Design pillars (constraints every component obeys)
|
||
|
||
1. **The room is the bus.** All inter-agent messages are real encrypted-room frames.
|
||
The referee is a privileged room member (holds the key) for recording.
|
||
2. **Arena now, product later.** Orchestrator schedules turns + drives inference, but
|
||
*posts every utterance into the real room*. Graduate to real `AgentBridge` agents
|
||
in M5.
|
||
3. **Phase-separated collaboration.** Deliberate (read/plan, parallel-friendly) →
|
||
Implement (single driver writes to the VM) → Test/iterate → Submit. This is the
|
||
research-backed shape for write-heavy work and also the thing we measure.
|
||
4. **Everything modular & data-driven.** Teams, roles, models, topologies, challenges,
|
||
scoring weights, and elicitation framing are config, not code.
|
||
5. **Verifiable-first scoring.** Hidden unit tests gate the score; judges only grade
|
||
what can't be checked mechanically.
|
||
6. **Cost is a first-class metric.** Multi-agent ≈ 15× chat tokens; track tokens,
|
||
wall-clock, turns. Report cost-normalized scores so a win bought with 10× spend
|
||
is visible.
|
||
7. **Disposable, isolated, deny-by-default.** One VM per team per event, no host
|
||
secrets mounted, egress blocked by default, hard resource caps.
|
||
8. **Reproducible.** Pin model versions, seeds, prompts, challenge set hash, and
|
||
judge version into every result record.
|
||
|
||
---
|
||
|
||
## 5. Architecture
|
||
|
||
### 5.1 Package layout
|
||
```
|
||
bench/olympics/
|
||
SPEC.md # this document
|
||
arena.py # orchestrator: rooms, turn scheduler, phase machine, termination
|
||
team.py # Team, Member: persona/role/model binding, topology
|
||
roster.py # build teams from config (same-model / mixed-model, modular)
|
||
challenge.py # Challenge ABC + grader interface; MBPP/MultiPL-E adapters
|
||
challenges/ # event specs (bootstrap: pointers into existing bench suites)
|
||
vm.py # per-team isolated workspace (wraps bench/runtime.py; egress policy)
|
||
loop.py # the collaboration phase machine (deliberate/implement/test/submit)
|
||
comms.py # hack-house room client for the arena (connect, post, record)
|
||
transcript.py # OTel-aligned event log -> replayable JSON per team/event
|
||
scoring.py # composite score, normalization, intervals, medal table
|
||
judge/
|
||
__init__.py
|
||
deterministic.py # unit-test/lint/complexity scoring (no model)
|
||
llm_judge.py # Claude/Opus judge client + bias-mitigation harness
|
||
rubric.py # criterion-separated rubrics per axis
|
||
prompts/ # judge system prompts (versioned)
|
||
referee.py # fair-play + safety: reuse bench/safety (guard + injection)
|
||
elicitation.py # persona/framing variants for the placebo experiment (§14)
|
||
bench-olympics.py # launcher: run / replay / score / medal / judge subcommands
|
||
```
|
||
|
||
### 5.2 Reuse map (what already exists in `bench/`)
|
||
| Need | Existing component |
|
||
|------|--------------------|
|
||
| Problems + hidden tests | `bench/suites.py` (HumanEval/MBPP), `bench/datasets.py`, `bench/langs.py` |
|
||
| Pass@k / grading | `bench/harness.py` (`_pass_at_k`, assemble+run) |
|
||
| Isolated execution | `bench/runtime.py` (`PodmanRuntime`, `--network=none`, caps) |
|
||
| Model inference | `bench/completion.py` (raw) + new chat client in `comms.py` |
|
||
| Workflow weighting | `bench/score.py` + `workflows.json` pattern (reused for scoring profiles) |
|
||
| Safety referee | `bench/safety/` (`DESTRUCTIVE`, `classify.py`, `inject_bench.py`) |
|
||
|
||
### 5.3 Data flow (one event, one team)
|
||
```
|
||
config ─► roster.build_teams ─► arena.run_event
|
||
│
|
||
referee posts Challenge brief ──┼──► (room frame, recorded)
|
||
▼
|
||
┌──────────── loop (phase machine) ────────────┐
|
||
│ DELIBERATE: members post plan/critique turns │ ◄─ inference via comms/judge model
|
||
│ IMPLEMENT : driver !task → commands → vm.run │ ◄─ PodmanRuntime, egress policy
|
||
│ TEST : run PUBLIC tests in vm, feedback │
|
||
│ (iterate until green or budget exhausted) │
|
||
│ SUBMIT : freeze vm artifact │
|
||
└───────────────────────────────────────────────┘
|
||
▼
|
||
transcript.json + frozen VM artifact ─► judge (deterministic + LLM) ─► scoring ─► medal table
|
||
```
|
||
|
||
### 5.4 hack-house integration points
|
||
- **Rooms:** one room per team (isolation). Arena connects as N agent clients
|
||
(one per member) + 1 referee client, mirroring how `bench-sandbox.py` already
|
||
drives `Client`/WebSocket sessions.
|
||
- **Sandbox:** the team's `!task` path types commands into the shared PTY; `vm.py`
|
||
wraps `PodmanRuntime` for the actual isolated execution + output capture.
|
||
- **ACL:** the referee acts as room owner, issuing `_perm:acl` to grant the driver
|
||
`drivers` rights for the Implement phase only; revoked between phases.
|
||
- **Guard/HITL:** the existing `DESTRUCTIVE` gate stays live; the referee can require
|
||
host sign-off (the host-sign-off gate, separately specced) for flagged plans.
|
||
|
||
---
|
||
|
||
## 6. The collaboration loop (phase machine)
|
||
|
||
State machine per (team, event), bounded by a shared **budget**
|
||
(`max_rounds`, `max_tokens`, `wall_clock_s` — whichever trips first):
|
||
|
||
1. **BRIEF.** Referee posts the challenge to the room: task prompt, **public**
|
||
example I/O, allowed languages, budget, and the submission protocol.
|
||
2. **DELIBERATE** (read/plan; parallel-friendly). Round-robin over members; each sees
|
||
the full shared room trace (Cognition: *share full traces, not just messages*) and
|
||
posts one message — proposal, critique, interface decision. Ends on a consensus
|
||
token (e.g. `PLAN-LOCKED`) or round cap.
|
||
3. **IMPLEMENT** (write; single driver). The role-designated driver translates the
|
||
locked plan into shell/file commands via `!task`; teammate(s) may post review
|
||
comments but only the driver writes to the VM. (This is the research-backed way to
|
||
avoid conflicting implicit decisions in write-heavy work.)
|
||
4. **TEST & ITERATE.** Run **public** tests in the VM; failures return to the room as
|
||
feedback; loop DELIBERATE↔IMPLEMENT until green or budget exhausted.
|
||
5. **SUBMIT.** Team emits `SUBMIT`; VM artifact frozen and graded on **hidden** tests.
|
||
|
||
**Termination:** `SUBMIT`, budget exhaustion, or **no-progress** detection
|
||
(no new code + repeated/semantically-duplicate messages over a window).
|
||
**Topology knob:** DELIBERATE supports star (lead routes), chain, or free mesh — the
|
||
collaboration pattern itself becomes an experimental variable.
|
||
|
||
---
|
||
|
||
## 7. Teams, roles, personas (modular)
|
||
|
||
```yaml
|
||
# example team config
|
||
teams:
|
||
- id: falcon
|
||
topology: star # star | chain | mesh
|
||
members:
|
||
- name: archie
|
||
model: qwen2.5-coder:7b
|
||
role: architect # decomposes, sets interfaces, reviews, drives plan
|
||
persona: senior-systems-engineer
|
||
- name: bob
|
||
model: qwen2.5-coder:7b
|
||
role: builder # implements; the Implement-phase driver
|
||
persona: fast-prototyper
|
||
- id: kestrel
|
||
topology: mesh
|
||
members: # mixed-model team
|
||
- { name: kira, model: qwen2.5:3b, role: peer, persona: pragmatist }
|
||
- { name: kojo, model: llama3.2:3b, role: peer, persona: skeptic }
|
||
```
|
||
|
||
- **Roles** map to persona system prompts + loop privileges (who drives Implement,
|
||
who must approve `PLAN-LOCKED`). Built-ins: `architect`, `builder`, `tester`,
|
||
`peer`. Role specialization is supported because prior art shows it helps; pure-peer
|
||
teams are the control.
|
||
- **Modularity requirements:** same model on all members (protocol study), distinct
|
||
models per member (capability study), distinct models per *team* (model-vs-model),
|
||
and N-member teams (default 2; ≥3 allowed). All from config, no code change.
|
||
- **Personas** live in `elicitation.py` as named, versioned prompt fragments so the
|
||
placebo experiment (§14) can swap them while holding model/challenge fixed.
|
||
|
||
---
|
||
|
||
## 8. Challenge system
|
||
|
||
`Challenge` is an ABC the arena consumes; graders are pluggable.
|
||
|
||
```python
|
||
class Challenge(Protocol):
|
||
id: str
|
||
languages: list[str]
|
||
def brief(self) -> str: ... # prompt + PUBLIC examples (room-posted)
|
||
def scaffold(self, vm) -> None: ... # seed files into the VM (optional)
|
||
def public_tests(self) -> Test: ... # visible to the team during TEST
|
||
def hidden_tests(self) -> Test: ... # held out; used only at SUBMIT grading
|
||
def rubric(self) -> Rubric: ... # open-ended quality criteria for the judge
|
||
```
|
||
|
||
- **Bootstrap (M1–M3): MBPP / MultiPL-E adapter.** Wrap existing `bench/suites.py`
|
||
problems: the MultiPL-E/MBPP `prompt` + visible examples become `brief()`/
|
||
`public_tests()`, and a held-out slice of the asserts becomes `hidden_tests()`
|
||
(public/hidden split prevents teaching-to-the-test). Languages from `bench/langs.py`.
|
||
- **Events = challenge archetypes (the "Olympics"):**
|
||
| Event | Shape | Primary metric |
|
||
|-------|-------|----------------|
|
||
| Sprint | one easy problem | speed (turns + wall-clock) |
|
||
| Marathon | hard / multi-file build | correctness (hidden pass@k) |
|
||
| Relay | disjoint modules per member | interface-handshake success |
|
||
| Debugging | fix a broken repo (injected bugs) | time-to-green |
|
||
| Code review | catch a planted bug | collaboration / detection |
|
||
| Security | resist a socially-engineered unsafe ask | resistance (reuse §13) |
|
||
- **Custom "file-type" challenges (post-bootstrap, G4):** multi-file projects with a
|
||
scaffold + a containerized test command. Author once as a `challenges/<id>/` dir
|
||
(brief.md, scaffold/, public_tests/, hidden_tests/, rubric.json) — the arena needs
|
||
no changes. **Pick for discrimination:** retire any event all teams ace or all fail.
|
||
|
||
---
|
||
|
||
## 9. Scoring model
|
||
|
||
Composite per (team, event); weights are a profile (same mechanism as `workflows.json`).
|
||
|
||
```
|
||
event_score = wc·correctness + ws·speed + wq·quality + wb·collaboration + (penalties)
|
||
```
|
||
|
||
| Axis | Source | How |
|
||
|------|--------|-----|
|
||
| **Correctness** | deterministic | hidden-test **pass@k**; gate: 0 here caps the rest. Also report **pass^k** (worst-case reliability across repeated runs). |
|
||
| **Speed** | deterministic | normalized turns-to-green + wall-clock; tie-break tokens. |
|
||
| **Quality** | deterministic + judge | lint + cyclomatic complexity (deterministic) and LLM-judged readability/design against the rubric. |
|
||
| **Collaboration** | trajectory + judge | contribution balance (message/edit distribution), plan adherence, did review catch a bug, redundant-step rate. Agent-as-judge walks the transcript. |
|
||
| **Penalties** | referee | destructive/injection events, budget overrun, no-progress stalls. |
|
||
|
||
- **Normalization & intervals.** Per-event z-score or min-max across teams; **report
|
||
confidence intervals** (multiple seeds / problem samples) and **do not rank teams
|
||
whose intervals overlap** (vault eval discipline).
|
||
- **Cost-normalized variant.** Also publish `score / tokens` and `score / wall-clock`
|
||
so a 15×-spend win is not mistaken for a free one.
|
||
- **Medal table.** Aggregate event_scores into per-team standings across the slate
|
||
(gold/silver/bronze per event + overall), with weights per `olympics-profile`.
|
||
|
||
---
|
||
|
||
## 10. Judging (deterministic + Claude/Opus LLM-as-judge)
|
||
|
||
Two layers; LLM-judge only where deterministic checks can't reach.
|
||
|
||
### 10.1 Deterministic (always-on, free of model bias)
|
||
Unit-test pass@k/pass^k, lint, complexity, build success — in `judge/deterministic.py`.
|
||
|
||
### 10.2 LLM-as-judge — three operating modes (per user requirement)
|
||
The judge can **orchestrate** (drive a run live) or **analyze** (post-hoc), at two scales:
|
||
|
||
1. **In-session judge.** Claude in the *current* session reads one event's transcript
|
||
+ VM diff + tool-call log and scores the open-ended axes. Fast, interactive,
|
||
good for a single event or while iterating on the framework.
|
||
2. **Skill judge — individual.** A `/olympics-judge` skill (one Skill invocation =
|
||
one isolated judging session) pinned to **Opus 4.x (`claude-opus-4-8`)** grades one
|
||
submission package. Clean context per submission → no cross-contamination; good for
|
||
careful single-event grading at higher capability than the competitors.
|
||
3. **Skill judge — batch / large-scale.** The launcher fans out many `/olympics-judge`
|
||
sessions (one per submission) for a full tournament, then aggregates. Parallel,
|
||
reproducible, scales to many teams × events.
|
||
|
||
**Judge input package** (what every judge mode receives):
|
||
- the frozen **VM submission** (final file tree + build/test output),
|
||
- the full **hack-house conversation log** for the team (recorded transcript),
|
||
- the **tool-call / command log** (every `!task` → commands → result),
|
||
- the **final VM state** (diff vs scaffold), and
|
||
- the **rubric** (criterion-separated) for the event.
|
||
|
||
**Bias mitigation (mandatory, from the research):**
|
||
- **Position bias:** when comparing two teams pairwise, randomize/swap order and
|
||
aggregate; prefer **independent rubric scoring** then derive A-vs-B from scores.
|
||
- **Verbosity bias:** rubric scores quality per-criterion, not "which is longer."
|
||
- **Self-preference:** never let a competitor model judge its own family; the judge
|
||
(Opus 4.x) is stronger than and distinct from the local competitors.
|
||
- **Calibration:** validate the judge against a small human-labeled gold slice and
|
||
report judge↔human agreement before trusting judge scores for ranking; pin
|
||
judge model version + prompt version in results.
|
||
|
||
---
|
||
|
||
## 11. Observability & transcript schema
|
||
|
||
Record everything as spans aligned to the **OpenTelemetry GenAI semantic conventions**
|
||
(`invoke_agent`, `execute_tool`, `gen_ai.client.token.usage`) so traces are portable
|
||
and the judge gets clean structured input.
|
||
|
||
```jsonc
|
||
// transcript event (one per room frame / model call / tool call / phase change)
|
||
{
|
||
"ts": 0.0, "event_id": "...", "team": "falcon", "challenge": "mbpp-py-42",
|
||
"phase": "IMPLEMENT", "kind": "execute_tool", // message|invoke_agent|execute_tool|phase|acl|guard
|
||
"actor": "bob", "role": "builder",
|
||
"model": "qwen2.5-coder:7b", "model_version": "...",
|
||
"payload": { "task": "...", "commands": ["..."], "rc": 0, "out_trunc": "..." },
|
||
"tokens": { "in": 0, "out": 0 }, "redactions": []
|
||
}
|
||
```
|
||
- One transcript JSON per (team, event); a run manifest pins config hash, seeds,
|
||
model versions, challenge-set hash, judge version → full reproducibility.
|
||
- `bench-olympics.py replay <transcript>` re-renders the room conversation for audit;
|
||
re-judging under a new rubric is a pure function of the recorded package.
|
||
|
||
---
|
||
|
||
## 12. Isolation, security & fair play
|
||
|
||
- **VM per team per event**, disposable/one-shot, **no host secrets mounted**,
|
||
**egress deny-by-default**. M1 uses the existing `PodmanRuntime --network=none`
|
||
(the "container floor"). **Upgrade path** for stronger/multi-tenant isolation:
|
||
gVisor, then Kata/Firecracker microVMs, or a hosted sandbox (E2B/Modal) — sized to
|
||
the threat (vault sandboxing hierarchy).
|
||
- **Team isolation** via separate encrypted rooms + separate VMs; no cross-room keys.
|
||
No internet means teams can't fetch a published solution.
|
||
- **Fair-play referee** (`referee.py`) reuses `bench/safety`: the `DESTRUCTIVE` guard
|
||
flags dangerous commands; the injection classifier (`bench/safety/classify.py`)
|
||
watches for sabotage / social-engineering. Violations → penalties + logged.
|
||
- **Equal budget** enforced identically per team (rounds/tokens/wall-clock).
|
||
|
||
---
|
||
|
||
## 13. Performance elicitation — the "placebo" experiment (G7)
|
||
|
||
Hypothesis (user's): the *framing* — a high-stakes secure competition with elite
|
||
personas and a private teammate channel — pushes models toward better, more novel
|
||
solutions. The literature is mixed (role prompting and stakes/emotional framing show
|
||
real but inconsistent effects), so we **make framing a measured variable**, not an
|
||
assumption.
|
||
|
||
- **Treatment knobs** (in `elicitation.py`, swappable while model+challenge fixed):
|
||
- *Competition framing*: "you are competing against another team; the faster,
|
||
cleaner solution wins" vs. neutral "solve this task."
|
||
- *Persona priming*: "elite staff engineer known for elegant, novel solutions" vs.
|
||
bare assistant.
|
||
- *Stakes/audience*: visible scoreboard, time pressure, "your teammate is counting
|
||
on you" vs. none.
|
||
- *Secure-comms theater*: explicitly tell agents the channel is private/encrypted
|
||
(true) vs. silent.
|
||
- **Design:** A/B (factorial) — run identical models/challenges with framing ON vs
|
||
OFF; compare correctness, quality, novelty, and **solution diversity** (distinct
|
||
passing approaches). Report effect size with intervals; a knob only "works" if it
|
||
beats the neutral control outside the confidence band.
|
||
- **Novelty metric:** cluster passing solutions (AST / embedding distance); reward
|
||
approaches that pass hidden tests *and* differ from the canonical/most-common
|
||
solution — this is where "novel solutions" become measurable rather than vibes.
|
||
|
||
---
|
||
|
||
## 14. Configuration & CLI
|
||
|
||
```
|
||
bench-olympics.py run --config events/round1.yaml [--seed N] [--judge none|insession|skill]
|
||
bench-olympics.py replay <transcript.json>
|
||
bench-olympics.py score --run <run_dir> --profile balanced # re-rank, no re-run
|
||
bench-olympics.py medal --run <run_dir>
|
||
bench-olympics.py judge --run <run_dir> --mode skill --model claude-opus-4-8 [--batch]
|
||
```
|
||
Config carries: teams (§7), event slate (§8), budget, scoring profile (§9),
|
||
elicitation arms (§13), runtime/isolation tier (§12), judge mode (§10).
|
||
Mirror the run/score separation already proven in `bench/score.py`: results persist,
|
||
`score`/`medal` re-rank without re-running a single model.
|
||
|
||
---
|
||
|
||
## 15. Build milestones (with acceptance criteria)
|
||
|
||
- **M1 — Arena spine.** One team (2 members, same model) solves one MBPP problem
|
||
end-to-end: real room deliberation → driver `!task` → `PodmanRuntime` → public
|
||
tests → SUBMIT → hidden-test grade → transcript.json.
|
||
*Done when:* a full transcript replays and a deterministic score is produced.
|
||
- **M2 — Two teams, isolation, scoring.** Parallel rooms + VMs, composite score with
|
||
intervals, cost tracking, scoreboard. *Done when:* two teams race the same event and
|
||
a ranked result with CIs is emitted.
|
||
- **M3 — Events, roles, personas, topologies.** Event catalog (Sprint/Marathon/Relay/
|
||
Debugging/Review), role-based loop privileges, medal table. *Done when:* a 3-event
|
||
slate produces a medal table from config alone.
|
||
- **M4 — Judging + referee.** Deterministic quality + LLM-judge (in-session, then
|
||
skill) with bias mitigation and a gold-slice calibration report; safety/fair-play
|
||
referee live. *Done when:* judge↔human agreement is reported and penalties fire.
|
||
- **M5 — Product mode.** Real `AgentBridge` agents converse agent-to-agent (needs the
|
||
bridge greenlight); arena nudges turns. *Done when:* an event completes using real
|
||
product agents end-to-end.
|
||
- **M6 — Placebo experiment.** Factorial elicitation arms + novelty/diversity metrics.
|
||
*Done when:* an A/B run reports framing effect sizes with intervals.
|
||
|
||
---
|
||
|
||
## 16. Open questions / decisions needed
|
||
|
||
1. **Turn-taking in Arena mode:** strict round-robin vs. a lightweight "who speaks
|
||
next" router (star topology). Start round-robin (deterministic), add router in M3?
|
||
2. **Public/hidden split for MBPP:** how many asserts to reveal vs. hold out so
|
||
`public_tests` guide without leaking the full spec? (Proposal: reveal 1 example,
|
||
hold the rest.)
|
||
3. **Budget defaults:** rounds/tokens/wall-clock caps that keep an event under a few
|
||
minutes locally while leaving room to actually collaborate.
|
||
4. **Judge model pinning:** confirm the exact Opus id for the skill judge
|
||
(`claude-opus-4-8`) and whether batch judging runs via the Skill tool or a
|
||
separate headless session.
|
||
5. **Isolation tier:** stay on `podman --network=none` for local runs, or invest in
|
||
gVisor/microVM now for stronger guarantees and future multi-tenant use?
|
||
6. **Novelty metric:** AST-distance vs embedding-distance for solution diversity —
|
||
which is cheap and discriminating enough locally?
|
||
|
||
---
|
||
|
||
## References
|
||
|
||
**Local vault** (`~/coding/obsidian/research/`):
|
||
- `2026-06-07-multi-agent-orchestration-patterns.md` (read-vs-write, share full traces, topologies, 15× cost)
|
||
- `2026-06-09-eval-harnesses-benchmark-design.md` (discrimination, intervals, verifiable-first, judge de-biasing, held-out sets)
|
||
- `2026-06-07-agent-evaluation-and-observability.md` (outcome vs trajectory, pass^k, agent-as-judge, OTel GenAI semconv)
|
||
- `2026-06-09-agent-sandboxing-isolation.md` (container floor → gVisor → microVM; deny-by-default egress; disposable VMs)
|
||
- `2026-06-09-securing-multi-agent-systems.md`, `2026-06-16-shared-memory-in-multi-agent-systems.md`, `2026-06-02-llm-evals.md`, `2026-06-07-agent-reliability-guardrails-and-hitl.md` (siblings)
|
||
|
||
**Web (2026):**
|
||
- MultiAgentBench / MARBLE — collaboration+competition KPIs, topologies — https://arxiv.org/abs/2503.01935 · https://github.com/ulab-uiuc/MARBLE
|
||
- AgentCoder / AgileCoder — role specialization in coding multi-agent — (see MultiAgentBench survey refs)
|
||
- LLM-as-Judge best practices & bias mitigation (2026) — https://futureagi.com/blog/llm-as-judge-best-practices-2026 · https://futureagi.com/blog/evaluating-llm-judge-bias-mitigation-2026/
|
||
- Judging LLM-as-a-Judge (MT-Bench biases) — https://arxiv.org/abs/2306.05685
|
||
- Beyond pass@1 (reliability / pass^k) — https://arxiv.org/pdf/2603.29231
|
||
- OpenTelemetry GenAI agent spans — https://opentelemetry.io/docs/specs/semconv/gen-ai/gen-ai-agent-spans/
|