How well it works
What we measured, how, and where it is weaker.
Memory tools are easy to claim and hard to check, so here is what we measured and what it does not show.
The measurement
We test retrieval: given a question about a long conversation, does Hivemind bring back the memories that contain the answer? The test is LoCoMo, a public set of long conversations with questions about them. We ran it against the live service, not a lab copy.
| Questions | 330, spread across all ten conversations |
| Found what the question needed | 78.2% |
| 95% interval | 73.4% to 82.3% |
| Memories sent per question | about 10 |
| Searches that failed | 0 |
A plain keyword search finds about 15 points less on the same questions.
Where it is weaker
| Question type | Found |
|---|---|
| Single fact | 82.2% |
| When something happened | 78.7% |
| Combining several facts | 73.0% |
| Open-ended | 57.1% |
Questions that need several conversations stitched together are the hardest, and the open-ended group is small, so its figure is rough.
What this does not show
- It measures whether the right memory was retrieved, not whether a model then answered correctly. Answer quality depends on the model you use.
- It is one benchmark on English conversations. Your projects will differ.
- A cheaper approach, sending the raw conversation turns, uses slightly fewer tokens per question. Hivemind wins on finding the answer, not on size.
Check it on your own memory
The retrieval inspector shows, for any question, exactly what a tool would be sent and what was left out. That is the test that matters for your own work.