Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
90 changes: 77 additions & 13 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,7 +5,7 @@ agents. Documents are chunked, embedded and searched server-side; this package
wraps the official `goodmem` Python SDK and exposes it to CAMEL both as a
toolkit and as a `BaseRetriever`.

**Version 0.2.1.** Verified against GoodMem server **v1.0.320**.
**Version 0.3.0.** Verified against GoodMem server **v1.0.320**.

> **Upgrading from 0.1.0.** 0.1.0 talked to GoodMem over hand-written HTTP and
> had defects that were invisible from its return values — a failed search
Expand Down Expand Up @@ -45,9 +45,9 @@ By default the model sees exactly two tools:
| `goodmem_search` | `query`, `top_k` |
| `goodmem_remember` | `text`, `metadata` |

Every operational setting — which spaces are readable, which reranker, whether
a threshold applies, whether files can be uploaded — is fixed by you at
construction time. The model cannot widen its own access, pick another space,
Every operational setting — which spaces are readable, which reranker, which
LLM answers from the results, whether a threshold applies, whether files can be
uploaded — is fixed by you at construction time. The model cannot widen its own access, pick another space,
or turn on indexing waits.

Opt in to more:
Expand All @@ -61,8 +61,8 @@ Opt in to more:

### Ids must be UUIDs

Every GoodMem id this package handles — the `space_ids` and `reranker_id` you
configure, and the `memory_id`, `space_id` and `embedder_id` a tool or method
Every GoodMem id this package handles — the `space_ids`, `reranker_id` and
`llm_id` you configure, and the `memory_id`, `space_id` and `embedder_id` a tool or method
takes — must be a UUID. Anything else raises `GoodMemIdError`, naming the
argument, **before any request is made**, because the GoodMem SDK puts ids
into request paths unescaped: `delete_memory("../spaces/<id>")` would
Expand All @@ -74,7 +74,8 @@ An empty string is not a UUID either: for no reranker, pass
`reranker_id=None` or leave it out. `reranker_id=""` meant "no reranker" in
0.2.0 and is now refused at construction, so
`reranker_id=os.getenv("GOODMEM_RERANKER_ID", "")` fails at startup — write
`os.getenv("GOODMEM_RERANKER_ID") or None`.
`os.getenv("GOODMEM_RERANKER_ID") or None`. `llm_id=""` is refused the same
way; for no LLM, pass `llm_id=None` or leave it out.

## Retrieval results

Expand All @@ -96,6 +97,7 @@ An empty string is not a UUID either: for no reranker, pass
"partial": False, # True when the server reported a problem
"statuses": [], # what it reported
"resultSetId": "...",
"abstractReply": "...", # only with llm_id -- see "LLM answers"
}
```

Expand Down Expand Up @@ -127,6 +129,56 @@ Those hits are `scoreKind: "vector"`, flipped like any vector score, and
`min_score` is not applied to them, so a reranker threshold cannot discard
them; `partial` is set and `statuses` carries both codes.

## LLM answers

GoodMem can run one of its configured LLMs over the chunks a search retrieved
and return a grounded answer beside them. This is **off by default** and
**set by you**, like `reranker_id`: pass the UUID of a GoodMem LLM as `llm_id`
when you construct the toolkit. The model never sees or chooses it —
`goodmem_search` still takes only `query` and `top_k`.

```python
from camel_goodmem import GoodMemRetriever, GoodMemToolkit

toolkit = GoodMemToolkit(space_ids=["<space-uuid>"], llm_id="<llm-uuid>")

result = toolkit.goodmem_search("What is the canary?")
result["abstractReply"] # "The canary is **ORYX-2290** ..."
result["results"] # the hits, exactly as without an LLM

rows = GoodMemRetriever(toolkit).query("What is the canary?")
rows[0]["extra_info"]["goodmem_abstract_reply"] # the same answer
```

Where the answer appears:

- **`goodmem_search`** (and so the tool result the model reads): the
`abstractReply` key, a string.
- **`GoodMemRetriever.query()`**: `goodmem_abstract_reply` in every row's
`extra_info`, since CAMEL's retriever returns a plain list of rows. It is
one answer for the whole retrieval, repeated on each row so it survives a
caller keeping only some of them.
- The key is present whenever `llm_id` is set, and absent otherwise.

The id is sent as `llm_id` in the retrieval's post-processor config, beside
`reranker_id` when both are set. It is a UUID like every other id: anything
else raises `GoodMemIdError` before a request is made.

An LLM does not rerank. Hits keep their scores, `scoreKind` and order
exactly as without it; combine it with `reranker_id` if you want reranking
as well.

**When the LLM fails**, the search does not. The server reports
`SUMMARIZATION_FAILED` — plus `NOT_FOUND` when no LLM has that id — and still
returns the hits. You get the hits, `partial: True`, both statuses in
`statuses`, a `warning`, and `abstractReply: None` (in the retriever,
`goodmem_partial: True`, `goodmem_statuses` and `goodmem_abstract_reply:
None`). Nothing is raised and no hit is dropped. Measured live: an LLM id
that does not exist gave `[NOT_FOUND, SUMMARIZATION_FAILED]` with the hit
kept; a provider out of credits gave `SUMMARIZATION_FAILED` carrying the
provider's `429`. A reranker configured beside a failing LLM keeps its
reranker scores.

## Metadata filters

Filters are expressions evaluated server-side, not SQL. You set them when you
Expand Down Expand Up @@ -203,8 +255,9 @@ rows = retriever.query("what did I store?", top_k=5)
`query()` returns CAMEL's retriever shape — `similarity score`, `content path`,
`metadata`, `extra_info`, `text` — with GoodMem specifics under `extra_info`
(`goodmem_chunk_id`, `goodmem_memory_id`, `goodmem_space_id`,
`goodmem_score_kind`, `goodmem_raw_score`, `goodmem_partial`, and
`goodmem_statuses` when degraded).
`goodmem_score_kind`, `goodmem_raw_score`, `goodmem_partial`,
`goodmem_statuses` when degraded, and `goodmem_abstract_reply` when the
toolkit has an `llm_id`).

## Bringing your own client

Expand All @@ -218,6 +271,17 @@ toolkit = GoodMemToolkit(client=Goodmem(base_url=..., api_key=...))
An injected client keeps its own server, credentials and TLS settings, and is
never closed by the toolkit.

## Changes in 0.3.0

New, opt-in: an LLM answer from the retrieved chunks. See
[LLM answers](#llm-answers).

| Was (0.2.1) | Now |
| --- | --- |
| No way to ask for GoodMem's LLM post-processing: `GoodMemToolkit(llm_id=...)` raised `TypeError: unexpected keyword argument 'llm_id'`, and no request carried one, so the `abstractReply` the result parser could read never arrived | `llm_id` constructor argument, checked as a UUID before any request and sent in the post-processor config; the answer is `abstractReply` on `goodmem_search` and `goodmem_abstract_reply` in the retriever's `extra_info`. Live with an OpenRouter `qwen/qwen3-8b` LLM: "The fixture canary is **ORYX-2290** ..." |
| Not reachable: no LLM could be requested | A failing LLM keeps the hits: `partial: True`, `statuses` `[NOT_FOUND, SUMMARIZATION_FAILED]` for an id that does not exist, `[SUMMARIZATION_FAILED]` for a provider `429`, `abstractReply: None`, never an exception |
| Not reachable | `goodmem_search(query, top_k)` is unchanged: the model cannot set or see the LLM |

## Changes in 0.2.1

Measured against a local server that records every request line, driving the
Expand Down Expand Up @@ -263,9 +327,9 @@ against GoodMem v1.0.320.

| Suite | Count | Needs |
| --- | --- | --- |
| `tests/test_goodmem_toolkit.py` | 86 | nothing — the real SDK over a mock transport, fed NDJSON captured from a live server |
| `tests/test_goodmem_ids.py` | 380 | nothing — the real SDK and `httpx` against a local server that records every request; every id-taking entry point (method, CAMEL tool, MCP tool, configuration) × ten malformed ids must send nothing, and a `str` or `uuid.UUID` subclass cannot change the id after it is checked. It also runs the live tests that depend on the id check against that server, and fails if any other live test passes an id the check would refuse |
| `tests/test_goodmem_live.py` | 30 | `GOODMEM_API_KEY` + `GOODMEM_BASE_URL`; skips entirely without them |
| `tests/test_goodmem_toolkit.py` | 101 | nothing — the real SDK over a mock transport, fed NDJSON captured from a live server |
| `tests/test_goodmem_ids.py` | 416 | nothing — the real SDK and `httpx` against a local server that records every request; every id-taking entry point (method, CAMEL tool, MCP tool, configuration) × ten malformed ids must send nothing, and a `str` or `uuid.UUID` subclass cannot change the id after it is checked. It also runs the live tests that depend on the id check against that server, and fails if any other live test passes an id the check would refuse |
| `tests/test_goodmem_live.py` | 35 | `GOODMEM_API_KEY` + `GOODMEM_BASE_URL`; skips entirely without them. The LLM tests also take `GOODMEM_TEST_LLM_ID` (a working LLM), and optionally `GOODMEM_TEST_FAILING_LLM_ID` (one whose provider fails) and `GOODMEM_TEST_RERANKER_ID`; each skips without its id |

```bash
pip install -e ".[dev]"
Expand All @@ -275,7 +339,7 @@ pytest tests/test_goodmem_toolkit.py tests/test_goodmem_ids.py

# live (pin the embedder if the server's first one is unhealthy)
GOODMEM_API_KEY=... GOODMEM_BASE_URL=... \
GOODMEM_TEST_EMBEDDER_ID=... \
GOODMEM_TEST_EMBEDDER_ID=... GOODMEM_TEST_LLM_ID=... \
pytest tests/test_goodmem_live.py

# what CI runs
Expand Down
2 changes: 1 addition & 1 deletion camel_goodmem/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -20,7 +20,7 @@
from camel_goodmem.retriever import GoodMemRetriever
from camel_goodmem.toolkit import GoodMemError, GoodMemToolkit

__version__ = "0.2.1"
__version__ = "0.3.0"

__all__ = [
"GoodMemToolkit",
Expand Down
22 changes: 19 additions & 3 deletions camel_goodmem/retriever.py
Original file line number Diff line number Diff line change
Expand Up @@ -24,7 +24,8 @@ class GoodMemRetriever(BaseRetriever):
Args:
toolkit (Any): A configured
:class:`~camel_goodmem.GoodMemToolkit`, which carries the
connection, the spaces and any metadata filter.
connection, the spaces, any reranker or LLM, and any metadata
filter.
metadata_filter (Optional[Union[Dict[str, Any], str]]): A filter
every result must also match: a mapping (an ``AND`` of
equalities) or an expression built with
Expand Down Expand Up @@ -91,7 +92,10 @@ def query(
``similarity score``, ``content path``, ``metadata``,
``extra_info`` and ``text``. When the retrieval was degraded
and nothing usable came back, a single dictionary is returned
whose ``text`` states what the server reported.
whose ``text`` states what the server reported. When the
toolkit has an ``llm_id``, every row's ``extra_info`` carries
``goodmem_abstract_reply``: the LLM's answer, or ``None`` if
it failed.
"""
outcome = self.toolkit._retrieve(
query, top_k, narrow=resolve_filter(self.metadata_filter)
Expand All @@ -117,6 +121,16 @@ def query(
)
hits = kept

# One answer per retrieval, so it rides on every row: a CAMEL caller
# that keeps only the first row, or only rows above a threshold,
# still has it.
summary: dict[str, Any] = {}
if (
outcome.abstract_reply is not None
or getattr(self.toolkit, "llm_id", None) is not None
):
summary["goodmem_abstract_reply"] = outcome.abstract_reply

results: list[dict[str, Any]] = []
for hit in hits:
extra: dict[str, Any] = {
Expand All @@ -126,6 +140,7 @@ def query(
"goodmem_score_kind": hit.score_kind,
"goodmem_raw_score": hit.raw_score,
"goodmem_partial": outcome.partial,
**summary,
}
if outcome.partial:
extra["goodmem_statuses"] = outcome.status_dicts
Expand Down Expand Up @@ -159,6 +174,7 @@ def query(
"extra_info": {
"goodmem_partial": True,
"goodmem_statuses": outcome.status_dicts,
**summary,
},
}
]
Expand All @@ -168,7 +184,7 @@ def query(
f"No information relevant to {query!r} is stored in "
"the configured GoodMem space(s)."
),
"extra_info": {"goodmem_partial": False},
"extra_info": {"goodmem_partial": False, **summary},
}
]
return results
50 changes: 45 additions & 5 deletions camel_goodmem/toolkit.py
Original file line number Diff line number Diff line change
Expand Up @@ -31,6 +31,13 @@
"an empty string is refused rather than read as 'no reranker'."
)

#: The same refusal for ``llm_id``, so ``llm_id=os.getenv("X", "")`` fails at
#: startup rather than being sent as an id.
_NO_LLM_HINT = (
"To search without an LLM, pass llm_id=None or leave it out; "
"an empty string is refused rather than read as 'no LLM'."
)


class GoodMemError(RuntimeError):
r"""Raised when a GoodMem operation fails.
Expand Down Expand Up @@ -81,8 +88,8 @@ class GoodMemToolkit(BaseToolkit):
embedded and searched server-side. This toolkit wraps the official
``goodmem`` Python SDK and exposes a deliberately narrow set of tools to
the model -- a search and, optionally, a write -- while every operational
setting (which spaces, which reranker, whether uploads are possible) is
fixed by the developer at construction time.
setting (which spaces, which reranker, which LLM, whether uploads are
possible) is fixed by the developer at construction time.

Args:
base_url (Optional[str]): The base URL of the GoodMem server. Falls
Expand All @@ -105,6 +112,16 @@ class GoodMemToolkit(BaseToolkit):
retrieval. Without one, no relevance threshold is applied. Pass
``None`` for no reranker: an empty string is a malformed id and
is refused. (default: :obj:`None`)
llm_id (Optional[str]): The UUID of a GoodMem LLM to run over the
retrieved chunks. Its grounded answer is returned as
``abstractReply`` by ``goodmem_search`` and as
``goodmem_abstract_reply`` in the retriever's ``extra_info``.
Scores are unaffected: an LLM does not rerank. When the LLM
fails, the server reports ``SUMMARIZATION_FAILED`` (and
``NOT_FOUND`` for an LLM that does not exist); the hits are still
returned, ``partial`` is set and ``abstractReply`` is ``None``.
Pass ``None`` for no LLM: an empty string is a malformed id and
is refused. (default: :obj:`None`)
min_score (Optional[float]): Drop hits scoring below this value.
Applies only to reranker scores: not without ``reranker_id``, and
not when the server reports that reranking failed and returns
Expand Down Expand Up @@ -145,6 +162,7 @@ def __init__(
timeout: float | None = 30.0,
upload_dir: str | Path | None = None,
reranker_id: str | None = None,
llm_id: str | None = None,
min_score: float | None = None,
metadata_filter: dict[str, Any] | str | None = None,
allow_write: bool = True,
Expand Down Expand Up @@ -175,6 +193,11 @@ def __init__(
if reranker_id is not None
else None
)
self.llm_id = (
require_uuid(llm_id, "llm_id", hint=_NO_LLM_HINT)
if llm_id is not None
else None
)
self.min_score = min_score
# Resolved now so a bad filter fails at construction rather than on
# the first search; resolved again at use, as it is public.
Expand Down Expand Up @@ -281,6 +304,12 @@ def _require_reranker(self) -> str | None:
self.reranker_id, "reranker_id", hint=_NO_RERANKER_HINT
)

def _require_llm(self) -> str | None:
r"""Returns the configured LLM as a canonical UUID, if any."""
if self.llm_id is None:
return None
return require_uuid(self.llm_id, "llm_id", hint=_NO_LLM_HINT)

def _space_keys(self, narrow: str = "") -> list[dict[str, Any]]:
r"""Builds the ``spaceKeys`` payload, including any metadata filter.

Expand Down Expand Up @@ -309,9 +338,11 @@ def _retrieve(
toolkit's own. (default: ``""``)

Returns:
RetrievalOutcome: The hits and any statuses the server reported.
RetrievalOutcome: The hits, any statuses the server reported, and
the LLM's answer when one was configured and produced.
"""
reranker_id = self._require_reranker()
llm_id = self._require_llm()
kwargs: dict[str, Any] = {
"message": query,
"space_keys": self._space_keys(narrow),
Expand All @@ -320,6 +351,10 @@ def _retrieve(
}
if reranker_id:
kwargs["reranker_id"] = reranker_id
if llm_id:
# The SDK puts it in the post-processor config beside
# ``reranker_id``. It does not change the hits or their scores.
kwargs["llm_id"] = llm_id

try:
stream = self._client.memories.retrieve(**kwargs)
Expand Down Expand Up @@ -377,7 +412,9 @@ def goodmem_search(self, query: str, top_k: int = 5) -> dict[str, Any]:
chunks, each with its text and the metadata of the memory it
came from), ``partial`` (``True`` when the server reported a
problem during this search), ``statuses`` (what the server
reported), and ``query``.
reported), and ``query``. When the developer configured an
LLM, ``abstractReply`` holds its answer drawn from the
results, or ``None`` if it failed (see ``statuses``).
"""
outcome = self._retrieve(query, top_k)
result: dict[str, Any] = {
Expand All @@ -394,7 +431,10 @@ def goodmem_search(self, query: str, top_k: int = 5) -> dict[str, Any]:
# an empty result carries the reason rather than reading as a
# clean miss.
result["warning"] = outcome.warning_text()
if outcome.abstract_reply:
if outcome.abstract_reply is not None or self.llm_id is not None:
# Present whenever an LLM was asked for, so a failed one reads as
# None beside its SUMMARIZATION_FAILED status, not as a missing
# key.
result["abstractReply"] = outcome.abstract_reply
return result

Expand Down
2 changes: 1 addition & 1 deletion pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@ description = "GoodMem integration for CAMEL."
authors = [{ name = "PAIR Systems" }]
readme = "README.md"
requires-python = ">=3.10"
version = "0.2.1"
version = "0.3.0"
license-files = ["LICENSE"]
urls.homepage = "https://github.com/PAIR-Systems-Inc/goodmem_camel"
urls.source = "https://github.com/PAIR-Systems-Inc/goodmem_camel"
Expand Down
Loading
Loading