MCP Studies
Five measured studies of how MCP agents use stored and curated knowledge. The server under test is txn2/mcp-data-platform, Apache-2.0, also hosted as Plexara. The model stays fixed. One thing about the platform moves. Grading is deterministic. Every number recomputes from committed run data. The reports sit on Zenodo.
Five notes, in reading order: whether a knowledge layer makes an agent measurably better; when agents use stored knowledge; the instrument those runs sat on; knowledge pollution after a curation gate; and whether cross-references help agents complete documents.
A benchmark that can only agree with its author is marketing. This series is not a marketing page.

1
Does a Semantic Knowledge Layer Make an Agent Measurably Better?
A reproducible benchmark: 42.7% to 98.7% on knowledge-trap questions, and a platform that learns from empty
2026-07-22
2When Do Agents Use Stored Knowledge?
A benchmark that killed its own hypothesis: strong models re-derive what they can check, and a weak model trusts a stale note over the evidence in front of it
2026-07-26
3An Instrument for MCP Agent Studies
The harness behind two DOI-archived benchmark reports, and why a study premise can be killed in one working day
2026-08-01
4Knowledge Pollution: Verification Displacement, Capability, and the Price of a Curation Gate
A wrong note that passed review spread only where the agent could have checked it, and only on Haiku 4.5: 16 of 24, against 0 of 24 on Sonnet 5 and Opus 5
2026-08-17
5Do Cross-References Help LLM Agents Complete Documents?
The agent found every page certified unreachable. Links cut search cost and keep working when search is off; they do not change what it reaches
2026-08-20