I deleted my retrieval graph and let the agent grep
I deleted the retrieval graph from my knowledge system and kept a searchable text index. Claude Code had made me reconsider how much understanding belonged inside the indexer.
I was building a knowledge system around documents, mail, and other connected sources. It extracted entities, built relationships, generated embeddings, and tried to resolve different mentions into a shared identity. Before the agent could answer a question, I had already asked another system to decide what the documents meant.
Meanwhile, Claude Code was making a simpler approach credible: search the files, read the useful parts, follow a reference, search again. Its success mattered to me. The agent could build understanding while doing the work.
That became the direction for my indexer: make the sources a greppable space for the agent. In this system, that means indexed lexical search across prepared source text. The default retrieval path uses BM25, not a shell running grep over remote files.
The difficult part was taking that idea out of a local repository and making it work across remote sources, binary files, changing permissions, and multiple users.
What I took from Claude Code
Anthropic describes Claude Code’s approach as a combination of context supplied up front and information discovered when needed. Project instructions provide a starting point; tools such as glob and grep let the model explore and read selectively. The model can change its next search after seeing the last result. Anthropic’s account
That was the useful idea. I did not need every relationship to exist in a database before an agent could follow it. A reference in a document could lead to another search. A remembered decision could tell the agent which project, name, or date mattered.
There is a cost to this freedom. Each extra search can mean another tool call and model round. Anthropic makes that tradeoff explicit. Vercel’s knowledge-agent implementation shows a related design: synchronize sources into a filesystem snapshot, then let the agent explore it in a sandbox. Cognition’s retrieval work focuses on reducing the serial rounds that make exploration slow.
My problem sat between those approaches. I wanted the agent to direct the search, but I did not want it downloading and converting a remote collection every time someone asked a question.
Why I had built the graph
I started with a routing problem. The information already lived in documents, mail, and other systems. I wanted the agent to know where to look, then read the original source when it needed detail. I also wanted it to retain what it learned between visits. Finding a source and remembering an earlier decision were separate jobs.
The graph was my attempt to connect that landscape. A project might appear under one name in a document and another in an email. I wanted those mentions connected before the question arrived.
An early design confused keeping pointers to sources with learning only from titles and metadata. That was too little information. Avoiding a second authoritative copy of the source did not remove the need to read its contents. This distinction survived the graph: the index still needs searchable evidence from inside the document.
Embeddings helped find candidates that did not share exact wording. They also supported identity matching and, in the earlier page representation, combined visual and textual information. The graph added explicit relationships the agent could retrieve. These were useful capabilities to aim for.
The trouble was that reading a page correctly did not guarantee a useful graph. The extractor could give the same concept different categories or promote table headings into entities. Identity matching then had to decide whether those records represented different things or inconsistent descriptions of the same thing.
That added a second maintenance problem after extraction. Documents could be read independently, but their inferred identities were shared across the corpus. More workers could prepare more content; they could also produce more decisions for the shared identity layer to reconcile. Source updates meant maintaining both the text and the interpretation built around it.
I was maintaining extracted facts, identity decisions, vectors, consolidation, and search projections alongside the source documents. In the recorded pre-cut snapshot, entity and relationship records accounted for about 83% of the search records. That tells me how much derived state the design carried. It does not mean those records were unused, or that they occupied the same proportion of disk or RAM.
I first considered changing how the graph was built. Put the full text into the search engine, examine vocabulary and recurring phrases across the corpus, then infer only the relationships that mattered. That was an explored design, not the system I ultimately shipped. It still left me building and maintaining a model of the corpus.
The next decision was smaller: ingest for BM25, measure retrieval, and leave embeddings for a later backfill if the results called for them. The readable passages already had source references and access information of their own. They could remain useful after I removed the entity and relationship layer.
Interpret each document
- Read source
- Extract + resolve
- Graph + text
Shared identities and relationships must stay consistent as sources change.
Infer from the searchable corpus
- Read source
- Index full text
- Infer relationships
Use vocabulary and recurring phrases to decide which relationships matter.
Make evidence searchable
- Read source
- Text + references
- Filter + retrieve
Return passages within caller access. Defer embeddings until evaluation justifies them.
I wanted a retrieval path I could inspect and test without another model deciding relevance. That did not make embeddings inherently random; a fixed model and input can be reproducible. It meant fewer model-dependent representations to version, rebuild, and debug. The index could make a new document searchable without first settling how every name in it fitted the rest of the corpus.
A remote collection is not a local folder
A local code repository already gives an agent readable files, paths, and a cheap way to inspect them. My sources did not come with those properties.
A filename in a connector listing is metadata. A scanned PDF needs transcription. A workbook needs all its tabs and rows preserved. A remote read can involve network latency, conversion, or a rate limit. Access also has to remain specific to the caller.
I kept an indexing pipeline to pay those preparation costs when content changes. Connector synchronization discovers items and updates. Workers fetch eligible content, convert it into searchable text, attach source references and permissions, and deliver it to the search index. An unchanged, reusable extraction can avoid repeating expensive work when the same source item and version is encountered again.
Decode or convert
Readable textUse native text; transcribe when needed
Page textParse tabs, headers and rows
Summary + row groupsOne of the more instructive mistakes was in spreadsheets. A summary and a few sample rows made a table look represented without making its contents searchable. If the user asked about an identifier in an omitted row, no retrieval technique could find it. I changed the representation to preserve row groups, sheet names, and headers, within explicit limits. Acquisition also had to preserve every tab before extraction started.
The same distinction applied to pages. Native text could be reused; pages without usable text needed a model to read them. Removing graph extraction did not eliminate that model work. It stopped transcription from sharing its output budget with entity and relationship generation.
This preparation is what makes a remote source collection behave more like a useful working directory. The agent receives readable evidence without rebuilding that view during each conversation.
A completed transcription could still be wrong
Removing the graph did not solve extraction. When pages kept hitting the output limit, I questioned whether a cap would discard legitimate dense documents. The investigation found another cause: a model could finish reading the page and then repeat newlines until it exhausted the budget.
On a synthetic reproduction, the constrained output ran for 66.8 seconds and used 8,192 tokens before truncating. Removing the required structured wrapper let the same page finish in 115 tokens. That was a narrow reproduction, not a corpus-wide speedup.
A repetition penalty looked like another fix. On a dense synthetic control it stopped normally, but preserved only three of seven decimal numbers, compared with all seven without the penalty. A successful completion could put a wrong number into the searchable text. The penalty was rejected as a default; the wrapper was removed.
That decision needed revisiting. Later, seven selected failing pages still looped without the wrapper. Across two runs per page, the normal setting truncated 14/14 times; a different penalty setting truncated 0/14 times. It also changed wording on pages that already completed. The compromise was to use it only after a whole-page truncation, leaving the first attempt unchanged. Those calls show recovery from a loop, not verified transcription accuracy. The fidelity risk remains on that retry.
Normal completion, missing decimals
without the penalty
with the tested penalty
Decision: reject the penalty as a default. Removing the structured wrapper fixed the reproduced loop.
Recover only after a failed attempt
at the normal setting
with a tested penalty
- Normal first attempt
- Truncation only
- Penalized retry
Decision: keep the first attempt unchanged. Recovery improved on this selected sample; fidelity was not established.
The evaluation had failures of its own. One early run reported 100% success on scanned pages after testing zero scanned pages. Another prompt candidate raised non-empty answers from 75% to 100% on four pages tested twice, but the text-bearing pages had the same word recall. The apparent gain came from answering the near-blank page. A synthetic sparse control could not distinguish the prompts either.
That prompt change was reverted. The evaluation began checking both directions: recover text where it exists, and leave blank controls empty. Counting completions alone could reward invented content that would then become searchable evidence.
Greppable does not mean scanning everything on every turn
My default search uses BM25 over an inverted text index. It is not a shell command walking every remote file. The index applies access filters and returns bounded passages with source references. Phrase matching and conservative stemming help with ordinary language. I also built and evaluated literal matching for exact strings. That optional support remains in the search layer, but the current API does not enable it on the default route. The greppable-space idea is the design direction; the shipped default API is ranked text search, not a full remote grep interface.
I tested combining lexical ranking with literal boosts so that ordinary language and exact strings could be matched in the same search. The architectural goal was a searchable view across prepared sources, without repeated scans and remote reads.
A search can span the indexed sources the caller can access. The agent does not have to choose a provider, enumerate its folders, fetch several candidates, convert them, and only then discover which one contains the useful line. It can start with the matching passages and follow a source reference when more context is needed.
That creates an opportunity to reduce tool calls and input tokens. It is not yet a measured reduction for the entire agent. A difficult question can still require several searches, and returning a short passage helps only if it contains enough evidence to guide the next step.
Shared content does not mean shared access
Two users can have access to the same source document through separate connections. That does not make either user’s access transferable. Retrieval applies the caller’s access filters at search time and returns only passages within that scope.
This boundary is separate from extraction. An unchanged source item can sometimes reuse an existing extraction, but each reader still needs independently authorized search records. Saving conversion work does not grant permission to search the result. A remembered title or project name is a query hint, not proof of access.
- 01SOURCE CONTENT
An unchanged source item may have a reusable extraction.
- 02SEPARATE ACCESS
Each reader needs independently authorized search records.
- 03CALLER-SCOPED RETRIEVAL
Search returns only the passages that caller can access.
Memory supplies context for retrieval
Memory can supply the project, alias, or earlier decision that makes a vague query searchable across sources. In a synthetic example, “the rollout we postponed” becomes a search for Project Cedar and a supplier approval. That context can connect a search in mail with one in documents. It does not require a vector lookup: remembered meaning can guide a lexical query.
This is a partial replacement for the graph’s purpose. A remembered connection does not provide exhaustive graph traversal. Memory can be stale, miss an alias, or contain an incorrect fact. Important claims still need source verification, and memory does not grant access to a document the caller cannot read.
The simpler system still had to earn its results
The retrieval output changed from graph and vector results to readable passages with source references. The useful test was whether that simpler output still found the intended evidence.
A saved experiment compared the standard analyzer with conservative plural stemming. It used 189 known-item probes, with 12-word queries derived from source text and inflected to test wording changes. Recall@1 rose from 57.1% to 89.4%; recall@10 rose from 86.8% to 97.9%. The saved baseline and the evaluation code both survive in the repository history.
That is evidence for fixing a specific lexical mismatch. It is not evidence that 97.9% of user questions get correct answers, or that the new architecture matches the old graph on arbitrary questions.
Fixing plural mismatches
189 known-item probes · 12-word inflected queries
Recall@1. Recall@10: 86.8% → 97.9%. This measures known-item retrieval, not answer accuracy.Cheaper phrase matching
About 23% lower p90. 30 real queries, four repetitions; medians across repetitions.
Content-term phrases versus raw-term phrases. Companion known-item results differed by one probe.Less search-record state
Retaining text alone implies about 5.75× fewer records.
Record count, not disk bytes, RAM usage or a measured bill reduction. The old graph had consumers.A stricter term-match threshold then looked excellent on synthetic probes derived from document text. Those probes shared much of the wording of the material they were meant to find. Before reindexing, I asked whether we were removing context the agent needed to read in the actual document. That led to testing paraphrases.
The investigation recovered real queries, which were a better test than another synthetic perturbation. The proposed threshold returned nothing across the eligible real-query set. It had looked selective because it removed noise; it was removing the useful results too. The proposal was withdrawn before rollout.
That comparison tested whether previously retrieved candidates survived the filter, not whether a human had judged each question answerable. It was enough to reject the threshold. Good outputs on source-derived probes had concealed a failure on queries people actually asked. An empty result still cannot prove the information is absent.
Even the recovered query log needed scrutiny. Five of its initial 64 entries were tool-output payloads rather than user searches. Their punctuation triggered exact-match filters in a later experiment, creating an apparent regression. They were removed from the fixture. Recovering production traffic did not make it clean evaluation data.
The architectural decision was to expose term coverage as a signal alongside the results, instead of using the proposed stricter threshold to hide candidates. The agent could still read the evidence. That is a narrower claim than having search decide whether the corpus can answer a question.
Latency had similarly specific causes. Phrase boosts originally included common pairs such as “how to” and “of the.” Scanning their positions was expensive and added little discrimination. A recorded comparison on 30 real queries, repeated four times, changed the phrase construction to content terms and allowed for the intervening words. Query p90 fell from 240 ms to 184 ms, about 23%. In the accompanying known-item check, the raw-term and content-term variants differed by one probe.
This is a local search optimization, not a measurement of whole-answer latency. But it is the kind of improvement I wanted to be able to make: identify the mechanism, change it, and check both speed and retrieval.
Removing embeddings, then the model after search
There were two separate model dependencies to remove. Embeddings represented the query and documents as vectors. Reranking took an already retrieved set of passages and used another model to reorder them. Removing one did not remove the other.
The embedding path had coupled live search to bulk indexing. A backfill could occupy the embedding service while a question waited for its query vector. I first reserved capacity for interactive requests. Moving to lexical retrieval later removed that query-vector prerequisite entirely. The unused vector path was deleted, and the embedding service was eventually retired after its callers were gone. That removed a dependency and its maintenance cost; I do not have a matched measurement assigning a two-second saving to embedding removal alone.
The roughly two-second result came from reranking. On a sample of 60 documents, using 12-word queries taken from their text, the recorded comparison showed mean search time of 37 ms without reranking and 2,267 ms with it. Reranking improved recall@10 from 54% to 70%. It bought better retrieval on that sample for about 2.23 seconds more per search.
Before retrieval
Query → embedding model → vector search
Lexical retrieval removes the query-vector dependency. No isolated two-second saving is established for this change.
After retrieval
Candidate passages → reranking model → reordered results
Removing this stage avoids its measured delay, with a ranking tradeoff.
About 2.23 seconds of added search time
| Retrieval | Mean time | Recall@10 |
|---|---|---|
| Lexical | 37 ms | 54% |
| With reranking | 2,267 ms | 70% |
I chose to remove that stage from the search path. A question can need several searches, so a model cost paid on every retrieval compounds. The decision was to accept a ranking tradeoff for faster access to evidence, rather than keep tuning the model that imposed the delay.
The post-removal validation recorded ten search calls with a 38 ms median, ranging from 18 to 162 ms. The accompanying 13-task suite passed both runs of every task, as it had before removal. Only three tasks directly exercised retrieval of known present evidence; the others checked absence, general behavior, or direct answers. That is a small regression check, not proof of equal answer quality on arbitrary questions. The 38 ms median and the earlier 2,267 ms mean also come from different workloads, so they are not a paired speedup measurement.
The ranking experiment has limits too. Queries copied from source text favor lexical matching and can understate a reranker’s value on paraphrases. The defensible latency claim is the measured 2.23-second reranking cost in that comparison, not a two-second reduction in every complete answer.
The larger test also contradicted my preferred direction. I had asked for literal boosting to be built and measured against BM25. In the larger test, using source-derived 12-word queries, it left recall@10 at 75.0% while query p90 rose from 99 ms to 752 ms. There were three wins and three losses on whether the target document appeared in the top ten. The earlier small-sample gain had not held up.
That result does not reject exact search for every workload. It does explain why literal boosting remains off on the default route. A greppable knowledge space was the goal; that particular matching strategy had not earned its cost. Neither this test nor the earlier reranking experiment establishes that reranking is generally unnecessary.
Where the cost went
The new design removes graph extraction, identity resolution, retrieval embeddings, and graph consolidation from the indexing path. Live retrieval no longer waits for a query embedding or a reranking pass.
The remaining costs are easier to separate. Content still needs synchronization, conversion and sometimes transcription. Text storage and search still cost money. The rest of the application still incurs model costs. Additional searches and returned context can spend some of the tokens saved by removing work earlier in the pipeline.
- Entity and relationship extraction
- Identity resolution and graph consolidation
- Retrieval and identity embeddings
- Live query embedding and reranking
- Source synchronization and conversion
- Transcription where text is unavailable
- Text indexing, storage and search
- Model costs elsewhere in the application
The pre-cut record counts imply roughly 5.75× fewer search records if only the text records are retained. That is a footprint calculation, not a storage-bill measurement. I do not have a matched billing comparison that turns it into a percentage saving, or a paired end-to-end evaluation proving equal answer quality before and after graph removal.
There is evidence elsewhere that simpler tools can improve outcomes. Vercel’s file-based analytics agent reported fewer tokens and steps on a five-query comparison. There is evidence in the other direction too: Cursor’s same-model comparison found benefits from giving agents semantic search alongside grep. Neither result measures my system. They make the evaluation question more precise: which work does this agent need for these tasks, and what does it cost to perform it?
What became easier to maintain
When a source changes now, I need to refresh its searchable representation. I no longer have to revisit a set of inferred identities and relationships just to make the new text available.
Discovery, extraction, and search visibility are still separate states. Finding a file in a connector does not mean its pages were read. Finishing extraction does not mean the asynchronous update has reached the search index. I still need to track those handoffs and retry failed delivery.
The ingestion record tracks unreadable pages and rejects documents with too many failures, but a search result is not a report that every page was successfully read. Rebuilding a paged document can still leave it temporarily absent if old passages are removed before replacement extraction succeeds. Stale work must also be prevented from restoring deleted content.
Those are concrete failures I can inspect: a missed update, an unreadable page, an undelivered index update, or a bad search term. There is less inferred state to repair between the source and the retrieval result.
The tradeoff remains real. Lexical search can miss unfamiliar wording. Text transcription can lose visual distinctions. Memory does not recreate every graph relationship. Structured totals still belong in a governed query engine, not in a count of the first page of search hits.
Claude Code gave me confidence in the agent’s ability to navigate evidence. My indexer’s job became making that evidence available across the sources a server-side agent has to work with. Memory keeps the useful context between visits. The next layer I add will have to answer a question those two cannot handle, and earn its maintenance cost on that question.