Local search for agent memory — no vector DB, no LLM in the query path
fidx combines local keyword and embedding search without an LLM in the query path, reporting latency, retrieval quality, and limitations against QMD.
fidx is local hybrid search over files you want a human or agent to treat as durable memory. It combines SQLite FTS5/BM25 for exact names and identifiers with ONNX embeddings for fuzzy "that doc about X" queries, then fuses the two rankings with RRF. There is no vector database, retrieval service, API key, or LLM in the query path. The index is a local SQLite file you can copy, back up, or delete; after the embedding model downloads once, search is fully offline.
I built it because I wanted grep for meaning, but the LLM-hybrid local search setups I measured, including QMD's query mode, kept answering one question in ten seconds. The pattern was always the same: to squeeze out more recall, the query path grows an LLM — query expansion, reranking, sometimes both. On CPU that was 12–36 seconds per query. Fine for a human asking one question; fatal for an agent that wants to check memory several times per step, and honestly fatal for me too, because search you wait on is search you stop using.

fidx is my refusal of that trade. The only model work per query is a single embedding pass:
-
Hybrid retrieval. SQLite FTS5 BM25 catches exact names and identifiers; 768-dim ONNX embeddings catch "that doc that discussed the indexing project"; reciprocal-rank fusion (RRF, k=60) merges both rankings. A document found by both arms outranks a rank-1 single-arm hit — that's the whole hybrid bet.
-
No retrieval infrastructure. No Qdrant/Chroma/Postgres service, docker-compose, auth, or schema migrations. Documents, BM25, and vectors still live in a local SQLite index you can copy, back up, or delete.
-
CPU-only, offline. The embedding model downloads once; after that, no network calls, no API keys, no telemetry.
-
Built for agents.
fidx servekeeps the model and index hot on a unix socket;--jsongives structured output and per-query advice on whether to preserve recall, trim weak results, or use stricter calibrated abstention. Warm queries: 18–49 ms measured (single ONNX thread) on document corpora, 452 ms on a 92k-file code corpus.
The numbers, honestly

I benchmarked against QMD — the closest comparable local tool, and the project whose storage design fidx openly borrows from — on four corpora (20 Newsgroups docs, a synthetic WhatsApp-style chat corpus, and 92k files of real code across 10 languages), 500 known-item queries per corpus (150 on the small one).
Two things most search benchmarks skip. First, purity: a tool that always returns 10 results "buys" recall with a noisy tail, so I had an LLM judge label every returned document and scored noise-rate@10 and clean@10 alongside recall. Second, conceding where you lose.
With fidx's built-in adaptive truncation enabled (one flag — it ships off by default so callers can choose recall-first behavior), the useful strengths versus QMD's LLM-hybrid query mode are not just speed: fidx has lower noise on all four corpora, recall wins on docs/code, a docs-small recall tie, and comparable chat recall. The absolute latency gap is the important part: docs were R@10 0.962 vs 0.914, clean@10 0.482 vs 0.060, and 49 ms vs 36 s. Where QMD wins: its plain FTS mode is faster on the 92k code corpus (121 ms vs fidx's 452 ms brute-force vector scan), takes R@1 on two corpora, and has the cleanest raw FTS lists on docs/chat/code where measured. The full tables, per-language code results, and a threats-to-validity section that says exactly when not to trust these numbers (known-item bias, machine-generated queries, single machine) are in BENCHMARKS.md.
Why "no LLM in the query path" is the whole design
The 10-second latency of other tools isn't an implementation detail — it is the price of putting a generative model between you and the index. And it buys surprisingly little at personal-corpus scale: in my measurements QMD's LLM modes gained little or no recall over fidx's plain hybrid. Worse, the LLM is a reliability liability: in the per-repo benchmark, the only engine failures (3 of 10 repos DNF) were LLM-embed stalls and session expiries. fidx replaces the LLM stages with score-shape math: a deterministic tail-cut on the fused score curve, and fidx calibrate, which derives a per-corpus score floor from self-retrieval probes + gibberish negatives — no labels, no benchmark tuning, microseconds at query time.
Try it
uv tool install fmdidx # PyPI name; the command is `fidx`
fidx doctor
fidx collection add ~/notes --name notes
fidx index
fidx search "that doc that discussed the indexing timeline"
macOS needs Homebrew Python (Apple's sqlite lacks loadable extensions — fidx doctor will tell you). License details are in the repo: github.com/williamliu-ai/fidx.
If you run the benchmark and get different numbers, I want to know — that's what the harness is for.