# How memory search works: hybrid retrieval, excerpts and result budgets (/blog/how-memory-search-works)

![Illustration: A small friendly robot with a round camera-eye head looks through a magnifying glass at a long bookshelf and holds up one short torn page from a thick notebook.](/blog/how-memory-search-works/hero.webp)

Memory search in CoreSpeed is hybrid. One call to `memory__search_memory` runs exact and relaxed lexical matching, entity aliases and vector similarity, and fuses them into a single ranking. Long memories come back as the excerpt that matched. The whole result is bounded, with an `omitted` count when rows were dropped, and the top hit is always returned.

**Four signals, one ranking**

Signals, fused:

* Exact lexical match — a name, a ticket id, a tool name
* Relaxed lexical match — another word form, a near miss
* Entity aliases — one thing known by several names
* Vector similarity — the same idea in other words

All through **memory\_\_search\_memory** (one ranking, no separate knobs).

What comes back:

* The top hit, always — however large it is
* Whole up to 1,200 chars — longer: the passage that matched
* limit, default 10 — rows are dropped from the bottom
* omitted: N — how many rows were dropped

One call, four signals, one ranking, one bounded result.

## What does hybrid retrieval mean here? \[#what-does-hybrid-retrieval-mean-here]

A memory can be asked for in several ways, and no single signal covers them all. Four run together:

| Signal                | What it catches                                             |
| --------------------- | ----------------------------------------------------------- |
| exact lexical match   | a name, a ticket id, a code, a tool name typed the same way |
| relaxed lexical match | the same word in another form, or a near miss in spelling   |
| entity aliases        | a person, product or team known by more than one name       |
| vector similarity     | the same idea in different words                            |

The four scores are fused into one ranking, so the result is one ordered list, and there are no separate knobs to tune. You pass a query and get memories back. The LongMemEval post describes the production path in a little more detail: full-text search plus dense vectors, combined by reciprocal-rank fusion, all under row-level security.

The reason for the fusion is practical. A pure vector search is weak on exact strings, and a saved memory is often an exact string: a ticket id, a hostname, a person's handle. A pure lexical search is weak on paraphrase, and the agent asking the question rarely uses the words the memory was saved in. Running both and fusing them is what makes a two-word query from a fresh session land on the right memory.

## How are long memories returned? \[#how-are-long-memories-returned]

`memory__search_memory` and `memory__list_memory` return many memories at once, so each row is bounded. A memory of 1,200 characters or fewer is returned whole, and that is most of them. A longer memory comes back as an excerpt, marked in the result:

```json
{
  "id": "0b5f…",
  "memory": "Escalation runbook: page the on-call rotation first…",
  "truncated": true,
  "full_length": 5310
}
```

For a search hit, the excerpt is the passage that matched the query, the same chunk the ranking scored, rather than the head of the memory. `list_memory` has no query to be relevant to, so it excerpts from the start. Pass the `id` to `memory__get_memory` to read the whole memory. That read follows normal visibility: your own private memories plus the organization's shared ones.

This is the detail that keeps a runbook useful as a memory. The agent sees the paragraph that answered its question, and the id to fetch the rest.

## How is the whole result bounded? \[#how-is-the-whole-result-bounded]

Each row has a bound, and so does the result. If the memories that matched would together exceed what a client can receive, the lowest-ranked rows are dropped and the result says how many, as `"omitted": 29` next to the `memories` array.

The top hit is always returned, however large it is. The default `limit` of 10 does not reach the budget in practice; only an explicitly high `limit` does. When you see `omitted`, narrow the query or lower `limit` to see what was dropped.

The budget exists because a search result goes into a model's context. A result that overflows the client is worse than a shorter one that fits, and a shorter one that keeps the best match is better still. Hence the asymmetry: rows are dropped from the bottom, and the first row is never dropped.

## How should an agent query it? \[#how-should-an-agent-query-it]

In short queries, early. On every MCP initialize, the server sends instructions that tell the agent to search memory with two or three short queries at the start of a non-trivial task, and to save durable facts after significant work. The same instructions carry a short index of the caller's pinned memories, the ones marked to surface at the start of every session.

Two or three short noun-phrase queries beat one long sentence, because each query gives the lexical signals something exact to match while the vector signal handles the rest. Memory operations cost 0 credits per call, so the cost of a second query is a little latency; storage quotas and rate limits protect the shared service. The [memory page](/docs/memory) lists every tool, and the [MCP page](/docs/mcp) covers what the initialize instructions contain.

## Why does recall alone not finish the job? \[#why-does-recall-alone-not-finish-the-job]

Retrieval is the first boundary. The published LongMemEval-S run, 500 questions, held retrieval recall at 10 at 0.99 throughout, and still moved end-to-end accuracy from 80.0 to 94.2 with a flash-class reader. The gains came from the evidence budget, the reader model and the reader instructions. [That post](/blog/how-we-took-longmemeval-from-80-to-94-without-touching-retrieval) explains each lever, and [a second one](/blog/we-retrieved-the-memory-then-dropped-the-answer) looks at what happens after the right memory is in the prompt.

For an agent that means the search result is an input, and what the agent does with it still decides the outcome. The excerpt and budget rules above are there to make that input small, ranked and complete at the top. The engine behind the search is open source as Lore, at [https://github.com/corespeed-io/lore](https://github.com/corespeed-io/lore).

## FAQ \[#faq]

**Can I choose vector-only or lexical-only search?**
No. The signals are fused into one ranking behind `memory__search_memory`, with no separate knobs.

**Why did a search return only part of a memory?**
The memory is longer than 1,200 characters. The result carries `truncated: true` and `full_length`; read the whole memory with `memory__get_memory`.

**What does `omitted: 29` mean?**
The 29 lowest-ranked matches were dropped to keep the result within what the client can receive. Narrow the query or lower `limit`.

**Does a search cost credits?**
No. Memory operations cost 0 credits per call. Rate limits and storage quotas apply instead.

**Which memories does a search see?**
The caller's own private memories and the organization's shared ones. Identity comes from the credential, and the boundary is enforced in the database.