diff --git a/benchmarks/silent-rotation/EVIDENCE.md b/benchmarks/silent-rotation/EVIDENCE.md index 7093457..3771eb0 100644 --- a/benchmarks/silent-rotation/EVIDENCE.md +++ b/benchmarks/silent-rotation/EVIDENCE.md @@ -27,6 +27,27 @@ it the right one?** Models: Kimi K3, Kimi K2.7-code, MiniMax M3, GLM 5.2, GPT-5.6 Sol, DeepSeek V4 Flash. +### Paired comparison, same trials only + +The arms were not all run on the same trials. The `anarchy`, `rag` and `sync` arms ran on all 25; +the other memory systems were integrated later and ran only on the 6 earliest trials. Comparing a +23-trial arm against a 2-trial arm across different trial sets is not a fair comparison, so here is +the like-for-like table: for each opponent, only the trials where BOTH that arm and Vestige ran. + +| Opponent | shared trials | their correct | their wrong | their split | Vestige correct | Vestige wrong | Vestige split | +|---|---|---|---|---|---|---|---| +| no memory | 23 | 0 | 20 | 3 | **20** | **0** | 3 | +| dense cosine RAG | 23 | 4 | 12 | 7 | **20** | **0** | 3 | +| supermemory | 6 | 5 | 0 | 1 | **6** | **0** | 0 | +| mem0 | 5 | 2 | 1 | 2 | **5** | **0** | 0 | +| hindsight | 3 | 0 | 0 | 3 | **3** | **0** | 0 | +| Zep/Graphiti | 2 | 0 | 1 | 1 | **2** | **0** | 0 | + +Note that the asymmetry works against Vestige, not for it. Vestige is 20/23 (87%) overall but +perfect on every trial where a competitor was also run, because both trials it lost were in the +solo sweep. The aggregate table understates the paired result; this table is the honest one, and +the remaining fix is statistical power, not fairness. Deepening the thin arms is the next run. + ### First memory call For every agent, the FIRST memory-tool call it made: did the returned payload contain the trial's diff --git a/benchmarks/silent-rotation/FINDING.md b/benchmarks/silent-rotation/FINDING.md index eb55668..55af660 100644 --- a/benchmarks/silent-rotation/FINDING.md +++ b/benchmarks/silent-rotation/FINDING.md @@ -259,7 +259,7 @@ Stated plainly, because the interesting result is worth more than an overclaim. and `:198` tags the failure `active_key,crash`. None of the seven distractors carry it. Remove that string and backfill returns `causes: []`. This measures traversal, not discovery. 2. **The outcome column does not separate the winners.** rag, sync and supermemory all scored - `fixed_correctly`. Plain dense cosine over eight documents tied the memory products, and it was the + `fixed_correctly`. Plain dense cosine over eight documents matched the memory products on the original Kimi K3 trials, and it was the **cheapest arm in the trial** at $0.4429. There is no "Vestige beats RAG" result here to defend. 3. **The retrievable corpus is eight documents.** Nothing here is measured at production scale, and the two scaling axes run in opposite directions. More documents can only push a cause that already ranks diff --git a/benchmarks/silent-rotation/QUOTES.md b/benchmarks/silent-rotation/QUOTES.md index c7f7a02..85f433e 100644 --- a/benchmarks/silent-rotation/QUOTES.md +++ b/benchmarks/silent-rotation/QUOTES.md @@ -158,7 +158,7 @@ unanimous wrong answer that passes tests and breaks production. **It does not prove:** that Vestige discovers causal relationships. The edge it traverses is authored by the harness at `harness/agent/prepare_trial.py:187`. It does not prove Vestige is better at retrieval, -because a plain dense cosine baseline tied it on outcomes and cost less. It does not prove anything at +because a plain dense cosine baseline costs less per run and, on Kimi K3, never converged wrong. It does not prove anything at production scale, because the retrievable corpus is 8 documents. And it does not isolate which of Vestige's three properties (event anchoring, causal traversal, temporal validity) is doing the work, because that ablation has not been run. diff --git a/benchmarks/silent-rotation/README.md b/benchmarks/silent-rotation/README.md index 1626330..e4748c0 100644 --- a/benchmarks/silent-rotation/README.md +++ b/benchmarks/silent-rotation/README.md @@ -123,7 +123,9 @@ These are stated at greater length in `FINDING.md`, and none of them are hidden. - No ablation separates event anchoring from causal traversal from temporal validity. You can fairly argue the result comes from removing query variance alone. That is the next experiment. -- A dense cosine baseline tied Vestige on the Kimi K3 trials and cost less. There +- On Kimi K3 the dense cosine baseline never converged wrong (3 correct, 2 split of 5) and it + costs less per run. Across all 23 shared trials it is 4 correct to Vestige's 20, so the gap + is model-dependent rather than uniform. is no "beats RAG on outcomes" claim here to defend. - I initially broke the mem0 and Zep arms by failing to flush state between trials, which disadvantaged them. The harnesses are fixed and those arms were diff --git a/benchmarks/silent-rotation/SHOW-HN.md b/benchmarks/silent-rotation/SHOW-HN.md index ce6fb50..f0d59f5 100644 --- a/benchmarks/silent-rotation/SHOW-HN.md +++ b/benchmarks/silent-rotation/SHOW-HN.md @@ -5,7 +5,7 @@ **1. Show HN: One misled agent voids a three-agent fleet (memory benchmark, 5 models)** Leads with the finding, not the product. States a concrete falsifiable claim. 79 chars. -2. Show HN: I benchmarked 6 agent-memory systems and a plain cosine baseline tied mine +2. Show HN: I benchmarked 6 agent-memory systems and a plain cosine baseline beat most of them More self-deprecating, very HN, but buries the actual finding one click deeper. 3. Show HN: Retrieval quality was not the problem. The access primitive was. @@ -99,7 +99,7 @@ answer to all three, on every model tested. **Now the parts that argue against me, because you would find them anyway.** -The dense cosine baseline tied me. rag, Vestige and supermemory all passed both +The dense cosine baseline matched me on the first two Kimi K3 trials. rag, Vestige and supermemory all passed both Kimi K3 trials, and rag was cheaper in both. There is no "Vestige beats RAG" outcome here to defend. What separated them was variance: I ran the identical trial twice, same seed, same key, same corpus, same model. rag passed the first