Repository navigation
Conversation
…at a time EmbeddingStore.search() decoded every row with struct.unpack and scored it with _cosine_similarity(), three generator expressions over every component, so each semantic search was O(N*D) interpreted Python: 1.4 s at 20k x 384 and 31 s at 42k x 4096 on a synthetic index. Score each 500-row chunk with one matrix-vector product instead. Products are taken in float64, like the loop's Python floats, so scores agree to ~1e-16 and the ranking is the same: other dimensionalities and zero norms still score 0.0, and the stable sort keeps read order for equal scores. Without numpy (cloud-only installs) the original loop runs, kept as _search_pure_python(). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
numpy ships with the "embeddings" extra, which the test job does not install, so the vectorized path in EmbeddingStore.search() would only ever be skipped there. uv.lock reuses the numpy versions already locked for "embeddings". Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
With numpy, a search still reads every vector back from SQLite on every query: 2.3 s at 42k x 4096. A cache on the store would never be hit twice, because _embedding_search() opens and closes a store per query. CRG_VECTOR_CACHE=1 keeps the decoded vectors in a process-wide cache keyed by (database, provider, dims). An entry is valid while PRAGMA data_version on a long-lived watcher connection does not move. It is read before the rows, so a commit that races the load costs an extra reload, never a stale answer. Rows the searching connection has not committed bypass the cache, at most four entries are kept, evicting a database's last entry closes its watcher, and EmbeddingStore.clear_vector_cache() releases everything. The scores come from the same helpers as the uncached path, so the ranking does not change; the search tests run again with the cache on to pin it. With the matrix warm, a 42k x 4096 search takes 0.40 s. Off by default: the watcher is a read connection held for the life of the process, which on Windows keeps graph.db from being deleted or replaced, and the matrix costs vectors x dims x 4 bytes of RAM. Documented in the README's environment-variable table and its four translations. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
rnoap
force-pushed
the
perf/vectorized-embedding-search
branch
from
October 2, 2026 20:50
6fac4b1 to
2dc0668
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Pull Request
Linked issue
Closes #1089
What & why
EmbeddingStore.search()decoded every stored vector withstruct.unpackand scored it with_cosine_similarity(), three generator expressions over every component. That makes everysemantic_search_nodescall O(N·D) interpreted Python: about 1.5 s at 20k × 384 (the defaultmodel) and 35 s at 42k × 4096. Numbers are in the issue.
Three commits. The second and the third can each be dropped without touching the first:
perf(embeddings): score stored vectors with numpy, not one component at a time: the fix.test: install numpy with the dev extra so CI runs the numpy search path.perf(embeddings): opt-in cache of decoded vectors across searches: off unlessCRG_VECTOR_CACHE=1, described below.The first commit keeps the same query, the same 500-row
fetchmanychunks and the same bounded memory,but scores each chunk with one numpy matrix-vector product:
The stored components are float32, so every product is exact and only the summation order
differs (scores agree to ~1e-15). A vector whose dimensionality differs from the query's,
or whose norm is zero, still scores 0.0. A zero query still scores everything 0.0. The
final sort is stable, so equal scores keep read order, as
list.sort(reverse=True)did.embeddingsextra, next tosentence-transformers. Without it, for example on a cloud-only install,search()runsthe original loop, moved unchanged into
_search_pure_python().of 4 made
struct.unpackraise, and_embedding_search()then dropped the whole vectorside of the search. It now scores 0.0, like any other dimension mismatch.
numpyadded to thedevextra (pyproject.toml,uv.lock+4 lines) so thetestjob exercises the numpy path. Drop that commit if you would rather keep
devlean. Thenumpy tests then skip themselves with
importorskip.Third commit, off by default:
CRG_VECTOR_CACHE=1With the first commit, every search still reads all the vectors back from SQLite, which is
most of what is left at 42k × 4096. A cache on the store would never be hit twice, because
_embedding_search()opens and closes a store per query. So the third commit keeps thedecoded matrix in a process-wide cache instead, only when
CRG_VECTOR_CACHE=1:PRAGMA data_versionon a long-lived watcher connection does not move, and that value moves on every commit
from any other connection. The token is read before the rows, so a commit that races the
load costs one extra reload, never a stale answer. A new file identity gets a fresh
watcher. That covers
graph.dbreplaced on Linux or macOS; on Windows, the open watcherblocks the replacement itself.
entries are kept. Evicting a database's last entry closes its watcher, and
EmbeddingStore.clear_vector_cache()releases everything.search tests run a second time with the cache on, through a subclass, to pin that.
It is opt-in because of what it holds. The watcher is a read connection that stays open for
the life of the process, against
_get_store's "callers must close it" contract, and onWindows an open connection keeps
graph.dbfrom being deleted or replaced. The matrix alsocosts vectors × dims × 4 bytes of RAM, about 690 MB at 42k × 4096. The README's variable
table says both. If you don't want it, drop the commit; the first two stand alone. Nothing in
CI sets the variable outside its own test class, so the
windows-nativejob never opens awatcher.
Same machine, a new store per search as
_embedding_search()does, medians of 3. "Sameresults" compares every warm search with the uncached one:
The first search pays the full decode, roughly twice the matrix in memory while it is
decoded. The warm search still converts each 500-row slice to float64, so it ranks exactly
like the uncached path instead of trading precision for the last few hundred milliseconds.
bench_cache.py
Timing
Random float32 vectors in a temporary database, top-20, the median of 3 searches against one
run of the loop. Windows 11, Python 3.11, numpy 2.4.6. Script below.
What is left at 42k × 4096 is mostly reading 690 MB of blobs out of SQLite, which is the part
a cross-query cache would remove.
bench_search.py
How it was tested
The 107 failures are environmental. I ran the same 107 node IDs on an unmodified
stagingcheckout and all of them fail there too. This machine checks files out with CRLF
(
core.autocrlf=true). That breaks, for example, the bundled D3 asset's SRI check. None ofthe failing tests are in
tests/test_embeddings.py. The Linux jobs here are the real check.New tests in
tests/test_embeddings.py(TestEmbeddingStoreSearch):_search_pure_python(), withscores within 1e-12.
order. Another provider's rows are ignored.
limitapplies (including 0).[].sys.modules, the loop runs.TestEmbeddingStoreVectorCachesubclasses it, so the five tests above run again withCRG_VECTOR_CACHE=1, and it adds the cache's own:because Windows refuses the real replacement while the watcher is open.
clear_vector_cache()empties the cache and closes the watchers.Checklist
uv run pytest tests/ --tb=short -q. All oftest_embeddings.pypasses. The full suite has 107 failures on this Windows checkout, and the same
ones fail on
staging(see above)uv run ruff check code_review_graph/uv run mypy code_review_graph/ --ignore-missing-imports --no-strict-optionaldocs/, docstrings): docstrings, and a row forCRG_VECTOR_CACHEin the README's environment-variable table🤖 Generated with Claude Code