Memory for agents that run on their own
Keep an unattended or framework-built agent's state outside its context window, and clear instead of compacting.
An agent running in a script, a scheduler or CI has no one to tell it where it got to. Its context window fills, gets compacted, and the details that mattered go with it. The way out is to keep the state somewhere that is not the window.
Hivemind is that place. Anything that can make an HTTP call can use it, so there is no terminal, editor or plugin involved.
The pattern
Three moments, and only three:
- At the start of a run, ask once with the task as the question.
- At each natural break, a step finished or a decision made, save one short statement.
- When the window gets heavy, clear it and go back to step 1.
Between those, the agent works from its own recent messages. Anything older has already been saved, so letting it roll off the window costs nothing.
In code
This uses the TypeScript client. The same calls exist over the REST API and as MCP tools.
import { Hivemind } from "@get-hivemind/sdk";
const hm = new Hivemind({ apiKey: process.env.HIVEMIND_API_KEY! });
const project = "nightly-triage";
// Start of a run: ask once, with the task as the question.
async function startContext(task: string): Promise<string> {
try {
const { payload, memoryIds } = await hm.recall({ q: task, sourceRef: project, budget: 1500 });
return memoryIds.length > 0 ? payload : "";
} catch {
return ""; // a slightly less informed run, not a broken one
}
}
// A natural break: say how you know.
async function checkpoint(content: string, status: "verified" | "assumed") {
try {
await hm.remember({ content, kind: "decision", sourceRef: project, metadata: { verification: status } });
} catch {
// writes can stop when the monthly allowance is spent; reads still work
}
}A framework's step loop is the same shape. In LangGraph, for instance, the first function is what a first node would call and the second is what a node calls when it finishes a unit of work. We have run these two functions as a plain script against a real account, in a fresh process, with no terminal. We have not run them inside LangGraph itself, so treat that mapping as a sketch.
Three things that keep it honest
Save statements, not logs. "The quarantine list lives in
ci/quarantine.txt and three tests were added" is worth finding later. A
paragraph of tool output is not. See
what is worth remembering.
Say how you know. Marking a checkpoint verified or assumed means the
next run sees which of its notes were checked and which were taken on trust.
An agent that reports a step done when it skipped it looks the same as one that
did it, unless someone wrote the difference down. The mark is the writer's word,
not a check Hivemind ran.
Ask a question that can have no answer. An unrelated task comes back empty, and the example above treats empty as "carry on without". Do not pad it.
Clearing instead of compacting
A compaction is a summary, and the summary decides what survives. A clear
decides nothing, because the state is already saved. When your window passes a
threshold you choose, finish the step, save a checkpoint, clear, and call
startContext again with the current task.
The recall in step 1 costs a known amount: 1,500 tokens by default, and no
single memory is longer than 700 characters in it. On a small local model with
a small window, lower budget to fit.
Over MCP instead
If your harness speaks MCP, point it at the server in
Add persistent memory to any MCP client. The agent gets
recall, remember and resume_handoff and decides when to call them.
remember takes a status of verified or assumed. For an unattended run
the explicit calls above are easier to reason about: you choose the moments.