Tag: Benchmark
143
Do Cross-References Help LLM Agents Complete Documents?
The agent found every page certified unreachable. Links cut search cost and keep working when search is off; they do not change what it reaches
2026-08-20
142Knowledge Pollution: Verification Displacement, Capability, and the Price of a Curation Gate
A wrong note that passed review spread only where the agent could have checked it, and only on Haiku 4.5: 16 of 24, against 0 of 24 on Sonnet 5 and Opus 5
2026-08-17
135An Instrument for MCP Agent Studies
The harness behind two DOI-archived benchmark reports, and why a study premise can be killed in one working day
2026-08-01
131When Do Agents Use Stored Knowledge?
A benchmark that killed its own hypothesis: strong models re-derive what they can check, and a weak model trusts a stale note over the evidence in front of it
2026-07-26
128Does a Semantic Knowledge Layer Make an Agent Measurably Better?
A reproducible benchmark: 42.7% to 98.7% on knowledge-trap questions, and a platform that learns from empty
2026-07-22