Hivemind

How well it works

What we measured, how, and where it is weaker.

Memory tools are easy to claim and hard to check, so here is what we measured and what it does not show.

The measurement

We test retrieval: given a question about a long conversation, does Hivemind bring back the memories that contain the answer? The test is LoCoMo, a public set of long conversations with questions about them. We ran it against the live service, not a lab copy.

Questions330, spread across all ten conversations
Found what the question needed78.2%
95% interval73.4% to 82.3%
Memories sent per questionabout 10
Searches that failed0

A plain keyword search finds about 15 points less on the same questions.

Where it is weaker

Question typeFound
Single fact82.2%
When something happened78.7%
Combining several facts73.0%
Open-ended57.1%

Questions that need several conversations stitched together are the hardest, and the open-ended group is small, so its figure is rough.

What this does not show

  • It measures whether the right memory was retrieved, not whether a model then answered correctly. Answer quality depends on the model you use.
  • It is one benchmark on English conversations. Your projects will differ.
  • A cheaper approach, sending the raw conversation turns, uses slightly fewer tokens per question. Hivemind wins on finding the answer, not on size.

Check it on your own memory

The retrieval inspector shows, for any question, exactly what a tool would be sent and what was left out. That is the test that matters for your own work.