Skip to content

perf(embeddings): score stored vectors with numpy, not one component at a time - #1090

Open
rnoap wants to merge 3 commits into
tirth8205:stagingfrom
rnoap:perf/vectorized-embedding-search
Open

rnoap wants to merge 3 commits into
tirth8205:stagingfrom
rnoap:perf/vectorized-embedding-search

Conversation

@rnoap

@rnoap rnoap commented Oct 2, 2026 •

Copy link
Copy Markdown

Pull Request

Linked issue

Closes #1089

What & why

EmbeddingStore.search() decoded every stored vector with struct.unpack and scored it with
_cosine_similarity(), three generator expressions over every component. That makes every
semantic_search_nodes call O(N·D) interpreted Python: about 1.5 s at 20k × 384 (the default
model) and 35 s at 42k × 4096. Numbers are in the issue.

Three commits. The second and the third can each be dropped without touching the first:

  1. perf(embeddings): score stored vectors with numpy, not one component at a time: the fix.
  2. test: install numpy with the dev extra so CI runs the numpy search path.
  3. perf(embeddings): opt-in cache of decoded vectors across searches: off unless
    CRG_VECTOR_CACHE=1, described below.

The first commit keeps the same query, the same 500-row fetchmany chunks and the same bounded memory,
but scores each chunk with one numpy matrix-vector product:

  • Same ranking as the loop. Products are taken in float64, like the loop's Python floats.
    The stored components are float32, so every product is exact and only the summation order
    differs (scores agree to ~1e-15). A vector whose dimensionality differs from the query's,
    or whose norm is zero, still scores 0.0. A zero query still scores everything 0.0. The
    final sort is stable, so equal scores keep read order, as list.sort(reverse=True) did.
  • No new runtime dependency. numpy already ships with the embeddings extra, next to
    sentence-transformers. Without it, for example on a cloud-only install, search() runs
    the original loop, moved unchanged into _search_pure_python().
  • One behavioural difference, on corrupt data only. A blob whose size is not a multiple
    of 4 made struct.unpack raise, and _embedding_search() then dropped the whole vector
    side of the search. It now scores 0.0, like any other dimension mismatch.
  • numpy added to the dev extra (pyproject.toml, uv.lock +4 lines) so the test
    job exercises the numpy path. Drop that commit if you would rather keep dev lean. The
    numpy tests then skip themselves with importorskip.

Third commit, off by default: CRG_VECTOR_CACHE=1

With the first commit, every search still reads all the vectors back from SQLite, which is
most of what is left at 42k × 4096. A cache on the store would never be hit twice, because
_embedding_search() opens and closes a store per query. So the third commit keeps the
decoded matrix in a process-wide cache instead, only when CRG_VECTOR_CACHE=1:

  • Entries are keyed by (database, provider, dims). One is valid while PRAGMA data_version
    on a long-lived watcher connection does not move, and that value moves on every commit
    from any other connection. The token is read before the rows, so a commit that races the
    load costs one extra reload, never a stale answer. A new file identity gets a fresh
    watcher. That covers graph.db replaced on Linux or macOS; on Windows, the open watcher
    blocks the replacement itself.
  • Rows that the searching connection has not committed yet bypass the cache. At most four
    entries are kept. Evicting a database's last entry closes its watcher, and
    EmbeddingStore.clear_vector_cache() releases everything.
  • It is the same scoring code as the uncached path, so the ranking does not change. The
    search tests run a second time with the cache on, through a subclass, to pin that.

It is opt-in because of what it holds. The watcher is a read connection that stays open for
the life of the process, against _get_store's "callers must close it" contract, and on
Windows an open connection keeps graph.db from being deleted or replaced. The matrix also
costs vectors × dims × 4 bytes of RAM, about 690 MB at 42k × 4096. The README's variable
table says both. If you don't want it, drop the commit; the first two stand alone. Nothing in
CI sets the variable outside its own test class, so the windows-native job never opens a
watcher.

Same machine, a new store per search as _embedding_search() does, medians of 3. "Same
results" compares every warm search with the uncached one:

Vectors Dims numpy, no cache Cache, first search Cache, warm Same results
20,000 384 0.103 s 0.117 s 0.018 s yes
20,000 1536 0.426 s 0.494 s 0.079 s yes
20,000 4096 1.122 s 1.186 s 0.190 s yes
42,000 4096 2.298 s 3.994 s 0.401 s yes

The first search pays the full decode, roughly twice the matrix in memory while it is
decoded. The warm search still converts each 500-row slice to float64, so it ranks exactly
like the uncached path instead of trading precision for the last few hundred milliseconds.

bench_cache.py
import os
import statistics
import tempfile
import time
from pathlib import Path
from unittest.mock import patch

import numpy as np

from code_review_graph.embeddings import EmbeddingStore

SIZES = [(20_000, 384), (20_000, 1536), (20_000, 4096), (42_000, 4096)]
PROVIDER = "bench:provider"


class Provider:
    name = PROVIDER

    def __init__(self, query):
        self.query = query

    def embed_query(self, text):
        return self.query


def open_store(db, query):
    with patch("code_review_graph.embeddings.get_provider", return_value=Provider(query)):
        return EmbeddingStore(db)


def timed_search(db, query):
    store = open_store(db, query)
    try:
        start = time.perf_counter()
        result = store.search("q", limit=20)
        return time.perf_counter() - start, result
    finally:
        store.close()


def bench(n, dims, loops=3):
    rng = np.random.default_rng(0)
    query = rng.standard_normal(dims).astype(np.float32).tolist()
    with tempfile.TemporaryDirectory() as tmp:
        db = Path(tmp) / "graph.db"
        store = open_store(db, query)
        rows = (
            (f"f.py::n{i}", rng.standard_normal(dims).astype(np.float32).tobytes(), "h",
             PROVIDER)
            for i in range(n)
        )
        store._conn.execute("BEGIN")
        store._conn.executemany("INSERT INTO embeddings VALUES (?, ?, ?, ?)", rows)
        store._conn.execute("COMMIT")
        store.close()

        os.environ.pop("CRG_VECTOR_CACHE", None)
        timed_search(db, query)  # warm the page cache once
        uncached = statistics.median(timed_search(db, query)[0] for _ in range(loops))
        _, expected = timed_search(db, query)

        os.environ["CRG_VECTOR_CACHE"] = "1"
        try:
            first, _ = timed_search(db, query)
            warm_runs = [timed_search(db, query) for _ in range(loops)]
        finally:
            os.environ.pop("CRG_VECTOR_CACHE", None)
            EmbeddingStore.clear_vector_cache()
    warm = statistics.median(t for t, _ in warm_runs)
    same = all(result == expected for _, result in warm_runs)
    return uncached, first, warm, same


for n, dims in SIZES:
    uncached, first, warm, same = bench(n, dims)
    print(f"{n:,} x {dims}: {uncached:.3f} s -> first {first:.3f} s, warm {warm:.3f} s, "
          f"same results: {same}")

Timing

Random float32 vectors in a temporary database, top-20, the median of 3 searches against one
run of the loop. Windows 11, Python 3.11, numpy 2.4.6. Script below.

Vectors Dims Loop (before) numpy (after) Speed-up Same top-20 Max score delta
20,000 384 1.43 s 0.087 s 16x yes 3.3e-16
20,000 1536 5.54 s 0.386 s 14x yes 3.3e-16
20,000 4096 14.49 s 1.093 s 13x yes 3.0e-16
42,000 4096 31.36 s 2.116 s 15x yes 3.7e-16

What is left at 42k × 4096 is mostly reading 690 MB of blobs out of SQLite, which is the part
a cross-query cache would remove.

bench_search.py
import statistics
import tempfile
import time
from pathlib import Path
from unittest.mock import patch

import numpy as np

from code_review_graph.embeddings import EmbeddingStore

SIZES = [(20_000, 384), (20_000, 1536), (20_000, 4096), (42_000, 4096)]
PROVIDER = "bench:provider"


class Provider:
    name = PROVIDER

    def __init__(self, query):
        self.query = query

    def embed_query(self, text):
        return self.query


def bench(n, dims, loops=3):
    rng = np.random.default_rng(0)
    query = rng.standard_normal(dims).astype(np.float32).tolist()
    with tempfile.TemporaryDirectory() as tmp:
        with patch("code_review_graph.embeddings.get_provider", return_value=Provider(query)):
            store = EmbeddingStore(Path(tmp) / "graph.db")
        try:
            rows = (
                (f"f.py::n{i}", rng.standard_normal(dims).astype(np.float32).tobytes(), "h",
                 PROVIDER)
                for i in range(n)
            )
            store._conn.execute("BEGIN")
            store._conn.executemany("INSERT INTO embeddings VALUES (?, ?, ?, ?)", rows)
            store._conn.execute("COMMIT")
            store.search("q", limit=20)  # warm the page cache once

            fast = []
            for _ in range(loops):
                start = time.perf_counter()
                new = store.search("q", limit=20)
                fast.append(time.perf_counter() - start)

            start = time.perf_counter()
            old = store._search_pure_python(query, PROVIDER, 20)
            slow = time.perf_counter() - start
        finally:
            store.close()
    same = [name for name, _ in new] == [name for name, _ in old]
    delta = max(abs(a - b) for (_, a), (_, b) in zip(new, old))
    return slow, statistics.median(fast), same, delta


for n, dims in SIZES:
    slow, fast, same, delta = bench(n, dims)
    print(f"{n:,} x {dims}: {slow:.2f} s -> {fast:.3f} s ({slow / fast:.0f}x), "
          f"same top-20: {same}, max delta {delta:.1e}")

How it was tested

uv sync --extra dev
uv run pytest tests/test_embeddings.py -q   # 139 passed, 17 of them new
uv run pytest tests/ --tb=short -q -m "not browser and not upgrade"
                                            # 4162 passed, 820 skipped, 107 failed (see below)
uv run ruff check code_review_graph/        # All checks passed!
uv run --with types-networkx mypy code_review_graph/ --ignore-missing-imports --no-strict-optional --platform linux
                                            # Success: no issues found in 77 source files

The 107 failures are environmental. I ran the same 107 node IDs on an unmodified staging
checkout and all of them fail there too. This machine checks files out with CRLF
(core.autocrlf=true). That breaks, for example, the bundled D3 asset's SRI check. None of
the failing tests are in tests/test_embeddings.py. The Linux jobs here are the real check.

New tests in tests/test_embeddings.py (TestEmbeddingStoreSearch):

  • 1,234 random rows across three chunks rank exactly like _search_pure_python(), with
    scores within 1e-12.
  • Zero-norm vectors, other dimensionalities and orthogonal vectors score 0.0 and keep read
    order. Another provider's rows are ignored.
  • A zero query scores every row 0.0, and limit applies (including 0).
  • An empty index returns [].
  • With numpy blocked in sys.modules, the loop runs.

TestEmbeddingStoreVectorCache subclasses it, so the five tests above run again with
CRG_VECTOR_CACHE=1, and it adds the cache's own:

  • The cache is off without the variable: nothing is cached and no watcher is opened.
  • A second store reuses the same decoded entry.
  • A commit from another connection refreshes it.
  • Uncommitted rows bypass it and leave it untouched.
  • Evicting a database closes its watcher.
  • A new file identity replaces the watcher and the entry. The test simulates the new inode,
    because Windows refuses the real replacement while the watcher is open.
  • clear_vector_cache() empties the cache and closes the watchers.

Checklist

  • Tests added for new functionality
  • All tests pass: uv run pytest tests/ --tb=short -q. All of test_embeddings.py
    passes. The full suite has 107 failures on this Windows checkout, and the same
    ones fail on staging (see above)
  • Linting passes: uv run ruff check code_review_graph/
  • Type checking passes: uv run mypy code_review_graph/ --ignore-missing-imports --no-strict-optional
  • Lines are at most 100 characters
  • Docs updated where behavior changed (README, docs/, docstrings): docstrings, and a row for CRG_VECTOR_CACHE in the README's environment-variable table

🤖 Generated with Claude Code

rnoap and others added 3 commits October 2, 2026 13:30
…at a time

EmbeddingStore.search() decoded every row with struct.unpack and scored it
with _cosine_similarity(), three generator expressions over every component,
so each semantic search was O(N*D) interpreted Python: 1.4 s at 20k x 384
and 31 s at 42k x 4096 on a synthetic index.

Score each 500-row chunk with one matrix-vector product instead. Products
are taken in float64, like the loop's Python floats, so scores agree to
~1e-16 and the ranking is the same: other dimensionalities and zero norms
still score 0.0, and the stable sort keeps read order for equal scores.
Without numpy (cloud-only installs) the original loop runs, kept as
_search_pure_python().

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
numpy ships with the "embeddings" extra, which the test job does not
install, so the vectorized path in EmbeddingStore.search() would only ever
be skipped there. uv.lock reuses the numpy versions already locked for
"embeddings".

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
With numpy, a search still reads every vector back from SQLite on every
query: 2.3 s at 42k x 4096. A cache on the store would never be hit twice,
because _embedding_search() opens and closes a store per query.

CRG_VECTOR_CACHE=1 keeps the decoded vectors in a process-wide cache keyed
by (database, provider, dims). An entry is valid while PRAGMA data_version
on a long-lived watcher connection does not move. It is read before the
rows, so a commit that races the load costs an extra reload, never a stale
answer. Rows the searching connection has not committed bypass the cache,
at most four entries are kept, evicting a database's last entry closes its
watcher, and EmbeddingStore.clear_vector_cache() releases everything.

The scores come from the same helpers as the uncached path, so the ranking
does not change; the search tests run again with the cache on to pin it.
With the matrix warm, a 42k x 4096 search takes 0.40 s.

Off by default: the watcher is a read connection held for the life of the
process, which on Windows keeps graph.db from being deleted or replaced,
and the matrix costs vectors x dims x 4 bytes of RAM. Documented in the
README's environment-variable table and its four translations.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@rnoap
rnoap force-pushed the perf/vectorized-embedding-search branch from 6fac4b1 to 2dc0668 Compare October 2, 2026 20:50

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: EmbeddingStore.search() scores every vector in a pure-Python loop (1.4 s per search at 20k × 384, 31 s at 42k × 4096)

1 participant