Morrisons AI Solutions: intelligence today, a brighter tomorrow

Findings

We tested our own claims.

We ran Memory MCP against a public benchmark, exactly as we ship it, and we are publishing the numbers whether they flatter us or not.

What we ran

LongMemEval-S is a public benchmark of 500 questions, each about a set of roughly 48 realistic chat sessions the answer is hidden inside. We stored every session in Memory MCP with its real remember tool, asked each question through its real recall tool, and checked whether the right session came back in the top 5 results. Nothing was mocked: the real embedding model, the real database, the shipped default of content encrypted at rest.

The test ran in a disposable database created only for this benchmark. No customer or personal data was involved.

The result

98%found the right memory in the top 5, one memory per message

Stored the way the product is meant to be used — one distilled fact or message per memory, not a whole conversation dumped in one go — Memory MCP found the correct session in its top 5 results for 98.0% of the 500 questions (98.1% once the 30 unanswerable “trick” questions are set aside). Stored as whole sessions instead, which asks far more of the embedding model in one go, it still found 93.0%.

Top results checkedOne memory per messageOne memory per whole session
189.6%75.8%
396.0%89.8%
598.0%93.0%
1098.6%96.6%

By question type (top 5, one memory per message): recalling something the assistant said, 100%; a fact that changed over time, 100%; something spread across several sessions, 98.5%; reasoning about when something happened, 94.7%; a stated preference, 100%; something the user said in a single session, 98.6%.

What this does and does not show

Why we're telling you this

A small company selling a memory tool should be able to show its memory actually works. We would rather publish a real 98% than an unproven claim of something higher, and we will publish the next result too, better or worse.

Back to Memory MCP