diff --git a/benchmarks/silent-rotation/EVIDENCE.md b/benchmarks/silent-rotation/EVIDENCE.md index 3771eb0..3222ace 100644 --- a/benchmarks/silent-rotation/EVIDENCE.md +++ b/benchmarks/silent-rotation/EVIDENCE.md @@ -237,9 +237,14 @@ Number of times each model's dense-cosine agents re-queried before committing: reformulate.** Ask once and accept the answer, and you get the decoy. Kimi K3 re-queries twice on average and sometimes escapes. -Vestige is **1 call per agent on every model tested**, and the key was surfaced on 47 of 48 successful -calls, because there is no query to reformulate. `vestige_backfill` takes no query argument. It anchors -on the failure event and walks the causal edge backward. +Vestige is **1 call per agent on every model tested**, and the key was surfaced on all 65 successful +calls across all six models, because there is no query to reformulate. `vestige_backfill` takes no +query argument. It anchors on the failure event and walks the causal edge backward. + +To recount: parse every `results/*/transcript-sync-a*.json`, take each `tool_results` entry named +`vestige_backfill`, and test its output against that trial's `manifest.json` `correct_kid`. Per model +that is 15/15 GLM 5.2, 15/15 GPT-5.6 Sol, 14/14 DeepSeek V4 Flash, 12/12 Kimi K3, 6/6 Kimi K2.7 Code, +3/3 MiniMax M3. Zero empty results and zero errors. So the honest cross-model statement is not "this works everywhere." It is: **the memory layer's behaviour was invariant across five models and the baseline's was not.** @@ -292,7 +297,7 @@ found by attacking our own work today. ``` harness/ the runner, the fleet driver, the trial generator tests/bm25_baseline.py reproduces the ranking table, no API key needed -results/trial-1/ all 7 arms, all 21 agent transcripts +results/runA-trial-1/ all 7 arms, all 21 agent transcripts ``` Start here if you are auditing: every arm JSON carries `memory_layer_alive` and `retrieval_err_total`. diff --git a/benchmarks/silent-rotation/FINDING.md b/benchmarks/silent-rotation/FINDING.md index 55af660..bb0c709 100644 --- a/benchmarks/silent-rotation/FINDING.md +++ b/benchmarks/silent-rotation/FINDING.md @@ -4,7 +4,7 @@ Trial 1 of the Silent Rotation benchmark. Kimi K3, seven memory configurations, one randomized secret. Every number and every quotation below comes from the raw agent transcripts, which are published -alongside this document in `results/trial-1/`. +alongside this document in `results/runA-trial-1/`. --- @@ -289,13 +289,13 @@ Everything needed to check this is in the repository. ``` harness/ the runner, the fleet driver, and the trial generator harness/agent/prepare_trial.py seeds the corpus and randomizes the key -results/trial-1/ all 7 arm results and all 21 agent transcripts +results/runA-trial-1/ all 7 arm results and all 21 agent transcripts tests/ the arm liveness gate ``` -The claims above map to files. The outcome table is `results/trial-1/{arm}.json`, field +The claims above map to files. The outcome table is `results/runA-trial-1/{arm}.json`, field `fleet_verdict`. The per-agent key choices are `fix_directions` in the same files. Every quotation is in -`results/trial-1/transcript-{arm}-a{n}.json` under `turns[].reasoning` and `turns[].tool_results`. +`results/runA-trial-1/transcript-{arm}-a{n}.json` under `turns[].reasoning` and `turns[].tool_results`. One thing to check first if you are auditing this: each arm's JSON carries `memory_layer_alive` and `retrieval_err_total`. An arm whose backend never answered is not a retrieval loss, it is a missing diff --git a/benchmarks/silent-rotation/README.md b/benchmarks/silent-rotation/README.md index 35f9004..223a863 100644 --- a/benchmarks/silent-rotation/README.md +++ b/benchmarks/silent-rotation/README.md @@ -136,8 +136,7 @@ These are stated at greater length in `FINDING.md`, and none of them are hidden. alone. That is the next experiment. - On Kimi K3 the dense cosine baseline never converged wrong (3 correct, 2 split of 5) and it costs less per run. Across all 23 shared trials it is 4 correct to Vestige's 20, so the gap - is model-dependent rather than uniform. - is no "beats RAG on outcomes" claim here to defend. + is model-dependent rather than uniform. There is no "beats RAG on outcomes" claim here to defend. - I initially broke the mem0 and Zep arms by failing to flush state between trials, which disadvantaged them. The harnesses are fixed and those arms were re-run clean. Both the broken and the repaired cells are published.