How memory search works: hybrid retrieval, excerpts and result budgets
memory__search_memory fuses exact and relaxed lexical matching, entity aliases and vector similarity into one ranking, then bounds each row and the result.

Memory search in CoreSpeed is hybrid. One call to memory__search_memory runs exact and relaxed lexical matching, entity aliases and vector similarity, and fuses them into a single ranking. Long memories come back as the excerpt that matched. The whole result is bounded, with an omitted count when rows were dropped, and the top hit is always returned.
What does hybrid retrieval mean here?
A memory can be asked for in several ways, and no single signal covers them all. Four run together:
| Signal | What it catches |
|---|---|
| exact lexical match | a name, a ticket id, a code, a tool name typed the same way |
| relaxed lexical match | the same word in another form, or a near miss in spelling |
| entity aliases | a person, product or team known by more than one name |
| vector similarity | the same idea in different words |
The four scores are fused into one ranking, so the result is one ordered list, and there are no separate knobs to tune. You pass a query and get memories back. The LongMemEval post describes the production path in a little more detail: full-text search plus dense vectors, combined by reciprocal-rank fusion, all under row-level security.
The reason for the fusion is practical. A pure vector search is weak on exact strings, and a saved memory is often an exact string: a ticket id, a hostname, a person's handle. A pure lexical search is weak on paraphrase, and the agent asking the question rarely uses the words the memory was saved in. Running both and fusing them is what makes a two-word query from a fresh session land on the right memory.
How are long memories returned?
memory__search_memory and memory__list_memory return many memories at once, so each row is bounded. A memory of 1,200 characters or fewer is returned whole, and that is most of them. A longer memory comes back as an excerpt, marked in the result:
{
"id": "0b5f…",
"memory": "Escalation runbook: page the on-call rotation first…",
"truncated": true,
"full_length": 5310
}
For a search hit, the excerpt is the passage that matched the query, the same chunk the ranking scored, rather than the head of the memory. list_memory has no query to be relevant to, so it excerpts from the start. Pass the id to memory__get_memory to read the whole memory. That read follows normal visibility: your own private memories plus the organization's shared ones.
This is the detail that keeps a runbook useful as a memory. The agent sees the paragraph that answered its question, and the id to fetch the rest.
How is the whole result bounded?
Each row has a bound, and so does the result. If the memories that matched would together exceed what a client can receive, the lowest-ranked rows are dropped and the result says how many, as "omitted": 29 next to the memories array.
The top hit is always returned, however large it is. The default limit of 10 does not reach the budget in practice; only an explicitly high limit does. When you see omitted, narrow the query or lower limit to see what was dropped.
The budget exists because a search result goes into a model's context. A result that overflows the client is worse than a shorter one that fits, and a shorter one that keeps the best match is better still. Hence the asymmetry: rows are dropped from the bottom, and the first row is never dropped.
How should an agent query it?
In short queries, early. On every MCP initialize, the server sends instructions that tell the agent to search memory with two or three short queries at the start of a non-trivial task, and to save durable facts after significant work. The same instructions carry a short index of the caller's pinned memories, the ones marked to surface at the start of every session.
Two or three short noun-phrase queries beat one long sentence, because each query gives the lexical signals something exact to match while the vector signal handles the rest. Memory operations cost 0 credits per call, so the cost of a second query is a little latency; storage quotas and rate limits protect the shared service. The memory page lists every tool, and the MCP page covers what the initialize instructions contain.
Why does recall alone not finish the job?
Retrieval is the first boundary. The published LongMemEval-S run, 500 questions, held retrieval recall at 10 at 0.99 throughout, and still moved end-to-end accuracy from 80.0 to 94.2 with a flash-class reader. The gains came from the evidence budget, the reader model and the reader instructions. That post explains each lever, and a second one looks at what happens after the right memory is in the prompt.
For an agent that means the search result is an input, and what the agent does with it still decides the outcome. The excerpt and budget rules above are there to make that input small, ranked and complete at the top. The engine behind the search is open source as Lore, at https://github.com/corespeed-io/lore.
FAQ
Can I choose vector-only or lexical-only search?
No. The signals are fused into one ranking behind memory__search_memory, with no separate knobs.
Why did a search return only part of a memory?
The memory is longer than 1,200 characters. The result carries truncated: true and full_length; read the whole memory with memory__get_memory.
What does omitted: 29 mean?
The 29 lowest-ranked matches were dropped to keep the result within what the client can receive. Narrow the query or lower limit.
Does a search cost credits? No. Memory operations cost 0 credits per call. Rate limits and storage quotas apply instead.
Which memories does a search see? The caller's own private memories and the organization's shared ones. Identity comes from the credential, and the boundary is enforced in the database.