I deleted my retrieval graph and let the agent grep

indexingsearcharchitectureevaluation

I deleted the retrieval graph from my knowledge system and kept a searchable text index. Claude Code had made me reconsider how much understanding belonged inside the indexer.

INDEXING AND RETRIEVALFrom connected sources to an evidence-based response
01 BACKGROUND INDEXINGPrepare changed content and deliver searchable records asynchronously.02 AGENT SEARCH AND SOURCE READING queryquerycandidatespassages + referencesinspectrefine searchif neededsource contentcompose response SOURCE COLLECTIONConnected knowledgeOriginal content stays hereSYNCHRONIZATIONDiscover changes + removalsCapture source accessCONTENT PREPARATIONFetch · convert · transcribePreserve text + source linksSEARCH INDEXSearchable passagesReferences + access rulesUSER QUESTIONIntent + conversationCaller identityMEMORY CONTEXTEarlier decisions + aliasesContext for the next queryAGENTChoose what to searchDirect the next stepACCESS-FILTERED SEARCHSearch within caller accessRank matching passagesEVIDENCE REVIEWInspect returned passagesDecide what else to readSOURCE READOpen referenced sourceEnforce caller accessRESPONSECompose from evidenceAttach source references 01 BACKGROUND INDEXING02 SEARCH AND READAgent directs the next step.Need more? Refine the search.Enough evidence? Respond. asynchronous deliveryquery + caller contextpassages + referencesopen source if needed SOURCE COLLECTIONConnected knowledgeOriginal content stays hereSYNCHRONIZATIONDiscover changes + removalsCapture source accessCONTENT PREPARATIONFetch · convert · transcribePreserve text + source linksSEARCH INDEXSearchable passagesReferences + access rulesUSER QUESTIONIntent + caller identityMEMORY CONTEXTEarlier decisions + aliasesAGENTChoose what to searchDirect the next stepACCESS-FILTERED SEARCHSearch within caller accessRank matching passagesEVIDENCE REVIEWInspect returned passagesDecide what else to readSOURCE READOpen referenced sourceEnforce caller accessRESPONSECompose from evidenceAttach source references
Conceptual architecture of the current approach. Background indexing prepares changed content and delivers searchable passages with source references and access rules. The agent searches within caller access, reviews evidence, opens sources when needed, and can refine the query before responding. Memory supplies context, never permission. Index delivery is asynchronous; a missing result does not prove the source contains no answer.

I was building a knowledge system around documents, mail, and other connected sources. It extracted entities, built relationships, generated embeddings, and tried to resolve different mentions into a shared identity. Before the agent could answer a question, I had already asked another system to decide what the documents meant.

Meanwhile, Claude Code was making a simpler approach credible: search the files, read the useful parts, follow a reference, search again. Its success mattered to me. The agent could build understanding while doing the work.

That became the direction for my indexer: make the sources a greppable space for the agent. In this system, that means indexed lexical search across prepared source text. The default retrieval path uses BM25, not a shell running grep over remote files.

The difficult part was taking that idea out of a local repository and making it work across remote sources, binary files, changing permissions, and multiple users.

What I took from Claude Code

Anthropic describes Claude Code’s approach as a combination of context supplied up front and information discovered when needed. Project instructions provide a starting point; tools such as glob and grep let the model explore and read selectively. The model can change its next search after seeing the last result. Anthropic’s account

That was the useful idea. I did not need every relationship to exist in a database before an agent could follow it. A reference in a document could lead to another search. A remembered decision could tell the agent which project, name, or date mattered.

There is a cost to this freedom. Each extra search can mean another tool call and model round. Anthropic makes that tradeoff explicit. Vercel’s knowledge-agent implementation shows a related design: synchronize sources into a filesystem snapshot, then let the agent explore it in a sandbox. Cognition’s retrieval work focuses on reducing the serial rounds that make exploration slow.

My problem sat between those approaches. I wanted the agent to direct the search, but I did not want it downloading and converting a remote collection every time someone asked a question.

Why I had built the graph

I started with a routing problem. The information already lived in documents, mail, and other systems. I wanted the agent to know where to look, then read the original source when it needed detail. I also wanted it to retain what it learned between visits. Finding a source and remembering an earlier decision were separate jobs.

The graph was my attempt to connect that landscape. A project might appear under one name in a document and another in an email. I wanted those mentions connected before the question arrived.

An early design confused keeping pointers to sources with learning only from titles and metadata. That was too little information. Avoiding a second authoritative copy of the source did not remove the need to read its contents. This distinction survived the graph: the index still needs searchable evidence from inside the document.

Embeddings helped find candidates that did not share exact wording. They also supported identity matching and, in the earlier page representation, combined visual and textual information. The graph added explicit relationships the agent could retrieve. These were useful capabilities to aim for.

The trouble was that reading a page correctly did not guarantee a useful graph. The extractor could give the same concept different categories or promote table headings into entities. Identity matching then had to decide whether those records represented different things or inconsistent descriptions of the same thing.

That added a second maintenance problem after extraction. Documents could be read independently, but their inferred identities were shared across the corpus. More workers could prepare more content; they could also produce more decisions for the shared identity layer to reconcile. Source updates meant maintaining both the text and the interpretation built around it.

I was maintaining extracted facts, identity decisions, vectors, consolidation, and search projections alongside the source documents. In the recorded pre-cut snapshot, entity and relationship records accounted for about 83% of the search records. That tells me how much derived state the design carried. It does not mean those records were unused, or that they occupied the same proportion of disk or RAM.

I first considered changing how the graph was built. Put the full text into the search engine, examine vocabulary and recurring phrases across the corpus, then infer only the relationships that mattered. That was an explored design, not the system I ultimately shipped. It still left me building and maintaining a model of the corpus.

The next decision was smaller: ingest for BM25, measure retrieval, and leave embeddings for a later backfill if the results called for them. The readable passages already had source references and access information of their own. They could remain useful after I removed the entity and relationship layer.

INDEXING AND RETRIEVALHow the indexing decision changed
EARLIER · BUILT

Interpret each document

  1. Read source
  2. Extract + resolve
  3. Graph + text

Shared identities and relationships must stay consistent as sources change.

CONSIDERED · NOT SHIPPED

Infer from the searchable corpus

  1. Read source
  2. Index full text
  3. Infer relationships

Use vocabulary and recurring phrases to decide which relationships matter.

SHIPPED · TEXT FIRST

Make evidence searchable

  1. Read source
  2. Text + references
  3. Filter + retrieve

Return passages within caller access. Defer embeddings until evaluation justifies them.

Three decisions, not three deployed systems. The middle design was explored. The shipped text-first index retained source references and permissions; graph construction stopped being a prerequisite for search.

I wanted a retrieval path I could inspect and test without another model deciding relevance. That did not make embeddings inherently random; a fixed model and input can be reproducible. It meant fewer model-dependent representations to version, rebuild, and debug. The index could make a new document searchable without first settling how every name in it fitted the rest of the corpus.

A remote collection is not a local folder

A local code repository already gives an agent readable files, paths, and a cheap way to inspect them. My sources did not come with those properties.

A filename in a connector listing is metadata. A scanned PDF needs transcription. A workbook needs all its tabs and rows preserved. A remote read can involve network latency, conversion, or a rate limit. Access also has to remain specific to the caller.

I kept an indexing pipeline to pay those preparation costs when content changes. Connector synchronization discovers items and updates. Workers fetch eligible content, convert it into searchable text, attach source references and permissions, and deliver it to the search index. An unchanged, reusable extraction can avoid repeating expensive work when the same source item and version is encountered again.

Choose the route before choosing the representation
Text and HTML

Decode or convert

Readable text
PDF, slides and images

Use native text; transcribe when needed

Page text
Workbooks and tables

Parse tabs, headers and rows

Summary + row groups
Native text, page transcription and table extraction preserve different kinds of information.

One of the more instructive mistakes was in spreadsheets. A summary and a few sample rows made a table look represented without making its contents searchable. If the user asked about an identifier in an omitted row, no retrieval technique could find it. I changed the representation to preserve row groups, sheet names, and headers, within explicit limits. Acquisition also had to preserve every tab before extraction started.

The same distinction applied to pages. Native text could be reused; pages without usable text needed a model to read them. Removing graph extraction did not eliminate that model work. It stopped transcription from sharing its output budget with entity and relationship generation.

This preparation is what makes a remote source collection behave more like a useful working directory. The agent receives readable evidence without rebuilding that view during each conversation.

A completed transcription could still be wrong

Removing the graph did not solve extraction. When pages kept hitting the output limit, I questioned whether a cap would discard legitimate dense documents. The investigation found another cause: a model could finish reading the page and then repeat newlines until it exhausted the budget.

On a synthetic reproduction, the constrained output ran for 66.8 seconds and used 8,192 tokens before truncating. Removing the required structured wrapper let the same page finish in 115 tokens. That was a narrow reproduction, not a corpus-wide speedup.

A repetition penalty looked like another fix. On a dense synthetic control it stopped normally, but preserved only three of seven decimal numbers, compared with all seven without the penalty. A successful completion could put a wrong number into the searchable text. The penalty was rejected as a default; the wrapper was removed.

That decision needed revisiting. Later, seven selected failing pages still looped without the wrapper. Across two runs per page, the normal setting truncated 14/14 times; a different penalty setting truncated 0/14 times. It also changed wording on pages that already completed. The compromise was to use it only after a whole-page truncation, leaving the first attempt unchanged. Those calls show recovery from a loop, not verified transcription accuracy. The fidelity risk remains on that retry.

INDEXING AND RETRIEVALFinishing is not the same as preserving the source
A FIX THAT DAMAGED THE TEXT

Normal completion, missing decimals

7 / 7Decimals preserved
without the penalty
3 / 7Decimals preserved
with the tested penalty

Decision: reject the penalty as a default. Removing the structured wrapper fixed the reproduced loop.

THE LATER, LIMITED COMPROMISE

Recover only after a failed attempt

14 / 14Calls truncated
at the normal setting
0 / 14Calls truncated
with a tested penalty
  1. Normal first attempt
  2. Truncation only
  3. Penalized retry

Decision: keep the first attempt unchanged. Recovery improved on this selected sample; fidelity was not established.

Separate historical probes with different penalty settings. The decimal check used one dense synthetic page. The retry check selected seven already-failing pages and ran each twice. Finishing without truncation does not establish text fidelity.

The evaluation had failures of its own. One early run reported 100% success on scanned pages after testing zero scanned pages. Another prompt candidate raised non-empty answers from 75% to 100% on four pages tested twice, but the text-bearing pages had the same word recall. The apparent gain came from answering the near-blank page. A synthetic sparse control could not distinguish the prompts either.

That prompt change was reverted. The evaluation began checking both directions: recover text where it exists, and leave blank controls empty. Counting completions alone could reward invented content that would then become searchable evidence.

Greppable does not mean scanning everything on every turn

My default search uses BM25 over an inverted text index. It is not a shell command walking every remote file. The index applies access filters and returns bounded passages with source references. Phrase matching and conservative stemming help with ordinary language. I also built and evaluated literal matching for exact strings. That optional support remains in the search layer, but the current API does not enable it on the default route. The greppable-space idea is the design direction; the shipped default API is ranked text search, not a full remote grep interface.

I tested combining lexical ranking with literal boosts so that ordinary language and exact strings could be matched in the same search. The architectural goal was a searchable view across prepared sources, without repeated scans and remote reads.

A search can span the indexed sources the caller can access. The agent does not have to choose a provider, enumerate its folders, fetch several candidates, convert them, and only then discover which one contains the useful line. It can start with the matching passages and follow a source reference when more context is needed.

That creates an opportunity to reduce tool calls and input tokens. It is not yet a measured reduction for the entire agent. A difficult question can still require several searches, and returning a short passage helps only if it contains enough evidence to guide the next step.

Shared content does not mean shared access

Two users can have access to the same source document through separate connections. That does not make either user’s access transferable. Retrieval applies the caller’s access filters at search time and returns only passages within that scope.

This boundary is separate from extraction. An unchanged source item can sometimes reuse an existing extraction, but each reader still needs independently authorized search records. Saving conversion work does not grant permission to search the result. A remembered title or project name is a query hint, not proof of access.

INDEXING AND RETRIEVALReuse content preparation, preserve separate access
  1. 01
    SOURCE CONTENT

    An unchanged source item may have a reusable extraction.

  2. 02
    SEPARATE ACCESS

    Each reader needs independently authorized search records.

  3. 03
    CALLER-SCOPED RETRIEVAL

    Search returns only the passages that caller can access.

Conceptual access boundary. Reusing an eligible extraction does not transfer permissions between readers. Search results remain scoped to the authenticated caller.

Memory supplies context for retrieval

Memory can supply the project, alias, or earlier decision that makes a vague query searchable across sources. In a synthetic example, “the rollout we postponed” becomes a search for Project Cedar and a supplier approval. That context can connect a search in mail with one in documents. It does not require a vector lookup: remembered meaning can guide a lexical query.

This is a partial replacement for the graph’s purpose. A remembered connection does not provide exhaustive graph traversal. Memory can be stale, miss an alias, or contain an incorrect fact. Important claims still need source verification, and memory does not grant access to a document the caller cannot read.

The simpler system still had to earn its results

The retrieval output changed from graph and vector results to readable passages with source references. The useful test was whether that simpler output still found the intended evidence.

A saved experiment compared the standard analyzer with conservative plural stemming. It used 189 known-item probes, with 12-word queries derived from source text and inflected to test wording changes. Recall@1 rose from 57.1% to 89.4%; recall@10 rose from 86.8% to 97.9%. The saved baseline and the evaluation code both survive in the repository history.

That is evidence for fixing a specific lexical mismatch. It is not evidence that 97.9% of user questions get correct answers, or that the new architecture matches the old graph on arbitrary questions.

INDEXING AND RETRIEVALThree measurements, three different claims
SAVED EVALUATION BASELINE

Fixing plural mismatches

189 known-item probes · 12-word inflected queries

Standard analyzer
57.1%
Conservative stemming
89.4%
Recall@1. Recall@10: 86.8% → 97.9%. This measures known-item retrieval, not answer accuracy.
RECORDED QUERY EXPERIMENT

Cheaper phrase matching

240ms p90 before→184ms p90 after

About 23% lower p90. 30 real queries, four repetitions; medians across repetitions.

Content-term phrases versus raw-term phrases. Companion known-item results differed by one probe.
DERIVED FROM PRE-CUT COMPOSITION

Less search-record state

Graph-derived
Text

Retaining text alone implies about 5.75× fewer records.

Record count, not disk bytes, RAM usage or a measured bill reduction. The old graph had consumers.
Historical measurements with different scopes. None is a paired graph-versus-current comparison of end-to-end answer quality or total spending.

A stricter term-match threshold then looked excellent on synthetic probes derived from document text. Those probes shared much of the wording of the material they were meant to find. Before reindexing, I asked whether we were removing context the agent needed to read in the actual document. That led to testing paraphrases.

The investigation recovered real queries, which were a better test than another synthetic perturbation. The proposed threshold returned nothing across the eligible real-query set. It had looked selective because it removed noise; it was removing the useful results too. The proposal was withdrawn before rollout.

That comparison tested whether previously retrieved candidates survived the filter, not whether a human had judged each question answerable. It was enough to reject the threshold. Good outputs on source-derived probes had concealed a failure on queries people actually asked. An empty result still cannot prove the information is absent.

Even the recovered query log needed scrutiny. Five of its initial 64 entries were tool-output payloads rather than user searches. Their punctuation triggered exact-match filters in a later experiment, creating an apparent regression. They were removed from the fixture. Recovering production traffic did not make it clean evaluation data.

The architectural decision was to expose term coverage as a signal alongside the results, instead of using the proposed stricter threshold to hide candidates. The agent could still read the evidence. That is a narrower claim than having search decide whether the corpus can answer a question.

Latency had similarly specific causes. Phrase boosts originally included common pairs such as “how to” and “of the.” Scanning their positions was expensive and added little discrimination. A recorded comparison on 30 real queries, repeated four times, changed the phrase construction to content terms and allowed for the intervening words. Query p90 fell from 240 ms to 184 ms, about 23%. In the accompanying known-item check, the raw-term and content-term variants differed by one probe.

This is a local search optimization, not a measurement of whole-answer latency. But it is the kind of improvement I wanted to be able to make: identify the mechanism, change it, and check both speed and retrieval.

There were two separate model dependencies to remove. Embeddings represented the query and documents as vectors. Reranking took an already retrieved set of passages and used another model to reorder them. Removing one did not remove the other.

The embedding path had coupled live search to bulk indexing. A backfill could occupy the embedding service while a question waited for its query vector. I first reserved capacity for interactive requests. Moving to lexical retrieval later removed that query-vector prerequisite entirely. The unused vector path was deleted, and the embedding service was eventually retired after its callers were gone. That removed a dependency and its maintenance cost; I do not have a matched measurement assigning a two-second saving to embedding removal alone.

The roughly two-second result came from reranking. On a sample of 60 documents, using 12-word queries taken from their text, the recorded comparison showed mean search time of 37 ms without reranking and 2,267 ms with it. Reranking improved recall@10 from 54% to 70%. It bought better retrieval on that sample for about 2.23 seconds more per search.

INDEXING AND RETRIEVALTwo model stages, two separate removals

Before retrieval

Query → embedding model → vector search

Lexical retrieval removes the query-vector dependency. No isolated two-second saving is established for this change.

After retrieval

Candidate passages → reranking model → reordered results

Removing this stage avoids its measured delay, with a ranking tradeoff.

SAME-SAMPLE RERANKING COMPARISON

About 2.23 seconds of added search time

RetrievalMean timeRecall@10
Lexical37 ms54%
With reranking2,267 ms70%
Faster retrieval gave up the recall gain measured on this sample. This is not an embedding-removal benchmark or evidence of answer-quality parity.
Recorded comparison on a 60-document sample with source-derived 12-word queries. The reranking arm adds candidate rescoring to lexical retrieval. These are mean search times, not whole-answer latency. Literal overlap can understate reranking’s value on paraphrases.

I chose to remove that stage from the search path. A question can need several searches, so a model cost paid on every retrieval compounds. The decision was to accept a ranking tradeoff for faster access to evidence, rather than keep tuning the model that imposed the delay.

The post-removal validation recorded ten search calls with a 38 ms median, ranging from 18 to 162 ms. The accompanying 13-task suite passed both runs of every task, as it had before removal. Only three tasks directly exercised retrieval of known present evidence; the others checked absence, general behavior, or direct answers. That is a small regression check, not proof of equal answer quality on arbitrary questions. The 38 ms median and the earlier 2,267 ms mean also come from different workloads, so they are not a paired speedup measurement.

The ranking experiment has limits too. Queries copied from source text favor lexical matching and can understate a reranker’s value on paraphrases. The defensible latency claim is the measured 2.23-second reranking cost in that comparison, not a two-second reduction in every complete answer.

The larger test also contradicted my preferred direction. I had asked for literal boosting to be built and measured against BM25. In the larger test, using source-derived 12-word queries, it left recall@10 at 75.0% while query p90 rose from 99 ms to 752 ms. There were three wins and three losses on whether the target document appeared in the top ten. The earlier small-sample gain had not held up.

That result does not reject exact search for every workload. It does explain why literal boosting remains off on the default route. A greppable knowledge space was the goal; that particular matching strategy had not earned its cost. Neither this test nor the earlier reranking experiment establishes that reranking is generally unnecessary.

Where the cost went

The new design removes graph extraction, identity resolution, retrieval embeddings, and graph consolidation from the indexing path. Live retrieval no longer waits for a query embedding or a reranking pass.

The remaining costs are easier to separate. Content still needs synchronization, conversion and sometimes transcription. Text storage and search still cost money. The rest of the application still incurs model costs. Additional searches and returned context can spend some of the tokens saved by removing work earlier in the pipeline.

INDEXING AND RETRIEVALRemoved work and remaining cost
REMOVED FROM THE PATH
  • Entity and relationship extraction
  • Identity resolution and graph consolidation
  • Retrieval and identity embeddings
  • Live query embedding and reranking
STILL PAID FOR
  • Source synchronization and conversion
  • Transcription where text is unavailable
  • Text indexing, storage and search
  • Model costs elsewhere in the application
Total cost also depends on repeated searches, returned tokens, model choice and source changes.
A cost map, not a proportional chart. Removing model-dependent work lowers those components; total savings depend on the work that remains and the agent’s search behavior.

The pre-cut record counts imply roughly 5.75× fewer search records if only the text records are retained. That is a footprint calculation, not a storage-bill measurement. I do not have a matched billing comparison that turns it into a percentage saving, or a paired end-to-end evaluation proving equal answer quality before and after graph removal.

There is evidence elsewhere that simpler tools can improve outcomes. Vercel’s file-based analytics agent reported fewer tokens and steps on a five-query comparison. There is evidence in the other direction too: Cursor’s same-model comparison found benefits from giving agents semantic search alongside grep. Neither result measures my system. They make the evaluation question more precise: which work does this agent need for these tasks, and what does it cost to perform it?

What became easier to maintain

When a source changes now, I need to refresh its searchable representation. I no longer have to revisit a set of inferred identities and relationships just to make the new text available.

Discovery, extraction, and search visibility are still separate states. Finding a file in a connector does not mean its pages were read. Finishing extraction does not mean the asynchronous update has reached the search index. I still need to track those handoffs and retry failed delivery.

The ingestion record tracks unreadable pages and rejects documents with too many failures, but a search result is not a report that every page was successfully read. Rebuilding a paged document can still leave it temporarily absent if old passages are removed before replacement extraction succeeds. Stale work must also be prevented from restoring deleted content.

Those are concrete failures I can inspect: a missed update, an unreadable page, an undelivered index update, or a bad search term. There is less inferred state to repair between the source and the retrieval result.

The tradeoff remains real. Lexical search can miss unfamiliar wording. Text transcription can lose visual distinctions. Memory does not recreate every graph relationship. Structured totals still belong in a governed query engine, not in a count of the first page of search hits.

Claude Code gave me confidence in the agent’s ability to navigate evidence. My indexer’s job became making that evidence available across the sources a server-side agent has to work with. Memory keeps the useful context between visits. The next layer I add will have to answer a question those two cannot handle, and earn its maintenance cost on that question.