Skip to content
Draft
Show file tree
Hide file tree
Changes from 13 commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions CHANGELOG.d/decision-window-constant-sql.md
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
Replaced the string-concatenated `IN (?, ?, …)` list in the decision-window reader with a single JSON-array parameter expanded by SQLite `json_each`, so the SQL text is constant. Behaviour is unchanged; this clears the two blocking Semgrep `sqlalchemy-execute-raw-query` findings that fail every PR on `main`.
1 change: 1 addition & 0 deletions CHANGELOG.d/multi-model-combination-authority.md
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
Removed the unreleased multi-model selection and fan-out surface. It may return only with released fast-mlsirm calibration and Fugu/Conductor/TRINITY-compatible allocation evidence.
12 changes: 7 additions & 5 deletions contextual_orchestrator/orchestrator.py
Original file line number Diff line number Diff line change
Expand Up @@ -5269,19 +5269,21 @@ def load_decision_window(self, limit: int = 256) -> dict[str, Any]:
phases = []
diagnostics = []
if request_ids:
placeholders = ",".join("?" for _ in request_ids)
# One JSON-array parameter expanded by SQLite's json_each keeps
# the statement text constant (no SQL string concatenation).
request_id_array = json.dumps(request_ids)
phases = self._conn.execute(
"SELECT kind, key, payload FROM orchestration_records "
"WHERE kind IN ('initial_decision', 'decision_receipt') "
"AND key IN (" + placeholders + ") "
"AND key IN (SELECT value FROM json_each(?)) "
"ORDER BY seq DESC LIMIT ?",
(*request_ids, 2 * limit + 1),
(request_id_array, 2 * limit + 1),
).fetchall()
diagnostics = self._conn.execute(
"SELECT kind, key, payload FROM orchestration_records "
"WHERE kind IN ('provider_dispatch', 'auxiliary_dispatch') "
"AND key IN (" + placeholders + ") ORDER BY seq DESC LIMIT ?",
(*request_ids, 8 * limit + 1),
"AND key IN (SELECT value FROM json_each(?)) ORDER BY seq DESC LIMIT ?",
(request_id_array, 8 * limit + 1),
).fetchall()
diagnostic_truncated = len(diagnostics) > 8 * limit
diagnostics = list(reversed(diagnostics[:8 * limit]))
Expand Down
65 changes: 65 additions & 0 deletions docs/doctoring/multi_model_combination.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,65 @@
# Multi-model combination: gap and reintroduction record

Status: Gap open (2026-09-27). A Proposed combination/fan-out surface was
added to PR #1267 and then withdrawn under review 5328347433, because it had
no selection authority. No combination code is present on this branch.

## Gap on `main` (`5665b0ad`)

The product's purpose is to raise answer quality by combining models (Sakana
Fugu; see [architecture](../architecture.md)). On `main`, no request path
combines two models' *answers*:

- `route_once` returns one worker's answer; judge rejection triggers
sequential failover to the next-ranked worker (`orchestrator.py:9171-9241`).
- `conduct` assigns exactly one worker per plan (`orchestrator.py:9467-9483`)
and the synthesizer only sees that single worker's output through the
access list.
- `nim_benchmark.evaluate_policies` compares route, conduct, and direct
single workers; it has no arm in which several workers answer the same task.

The claim that combining models beats the best single model is therefore
neither implemented nor measurable today.

## Why the Proposed surface was withdrawn

Review 5328347433 on PR #1267 found that the surface decided which model
answer to return without an authority for that decision:

- `RankedFirst` selected by caller-supplied rank, with no model.
- `PluralityVote` applied single-model self-consistency (Wang et al., 2023)
to different workers, whose errors correlate; no live paired evidence
existed.
- `ScoredBestOfN` accepted any float-returning callable, with no released
schema, calibration identity, uncertainty, or provenance.
- `collect_candidates` / `mixture_of_agents` accepted caller-chosen proposer
sets and concurrency without a Fugu/Conductor/TRINITY-compatible allocation
receipt.

Tie abstention (`0940e654`) repaired a symptom but not the missing
authority. Because the surface was not wired into serving, removal is the
fail-closed outcome.

## Reintroduction conditions

A combination surface may return only when all of these exist (from review
5328347433):

1. An immutable owner contract for the selection decision.
2. Released fast-mlsirm evidence for any scorer used to choose between
answers (schema, calibration identity, uncertainty, provenance).
3. A typed `no_decision` outcome instead of an implicit fallback.
4. Executable allocation and paired-evaluation acceptance: live paired arms
in `nim_benchmark` against `best_single_worker_hindsight`
(`paired_bootstrap_mean_difference`), and an allocation receipt for
proposer sets and concurrency.

Known prerequisites outside this record:

- Live evaluation needs the production provider keys, which exist only as
`production`-environment secrets usable from `main`; a `workflow_dispatch`
paid/free canary on `main` is required (see the coordination note on PR
#1263).
- ADR 0124 step 1 (PR #1262) and step 2 (ports/adapters) before serving
wiring.
- PR #1209 for the `main`-side test and security gates.
33 changes: 33 additions & 0 deletions docs/product-technical-gap-baseline.md
Original file line number Diff line number Diff line change
Expand Up @@ -7148,6 +7148,39 @@ suite: 3663 passed, 1 skipped, 5 known local-only failures (the openai SDK
`python -m interrogate -v contextual_orchestrator/` remains 100% (no production
code touched).

## 2026-09-27 Proposed multi-model selection authority repair (withdrawn) — PR #1267

**Observed exact-head gap.** PR #1267 head `a4183535` added a new decision
surface whose documentation explicitly described plurality `min_support` and
rank tie-breaking as repository choices rather than results of its cited
self-consistency algorithm. `ScoredBestOfN` also used rank when the scorer's
maximum tied. The fan-out API required every caller to supply a finite
wall-clock deadline, contradicting the default-null upstream-completion
boundary. These affect answer admission, model selection, and termination, so
they cannot remain uncalibrated policy controls.

**Action and evidence.** RED `e34a038b` records 12 focused failures for the
missing contracts. GREEN `0940e654` removes `min_support`, selects only a
unique plurality mode, abstains on equal maximum scores, and makes
`deadline_seconds=None` the default. A finite proposer deadline remains an
explicit administrative input; the completion port owns actual upstream
cancellation. Focused combination/fan-out tests are 58/58 GREEN and the four
changed Python files compile. Status remains **Proposed / PR-head only** until
exact-head hosted Security and Quality checks and an independent approval are
terminal GREEN; no merge or release claim follows from local evidence.

**Withdrawal (2026-09-27, review 5328347433).** Tie abstention repaired one
symptom but did not establish selection authority: `RankedFirst` selects by
caller rank, `PluralityVote` applies one-model self-consistency to different
workers whose errors correlate, `ScoredBestOfN` accepts an uncalibrated
callable, and the fan-out accepts caller-chosen proposer sets without an
allocation receipt. Because the surface was Proposed and not wired into
serving, it is removed from PR #1267 (fail closed). The combination gap and
its reintroduction conditions are recorded in
[multi-model combination](doctoring/multi_model_combination.md). PR #1267
keeps only the constant-SQL decision-window fix and the hash-locked
`fast-mlsirm`/`anyio`/Rust-toolchain CI repairs.

## 2026-09-14 Generated-plan step bound origin (section 3.1 / 5.1 fidelity)

`OrchestrationPolicy.max_workflow_steps = 6` bounded generated Conductor plans
Expand Down
2 changes: 1 addition & 1 deletion pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -11,7 +11,7 @@ dependencies = [
"opentelemetry-exporter-otlp-proto-http>=1.30.0",
"jsonschema>=4.26.0",
"egressweave>=0.1.0,<0.2.0",
"fast-mlsirm @ git+https://github.com/ContextualWisdomLab/fast-mlsirm.git@09f762ded35786dd1078222a4577ff09d649816f ; python_full_version >= '3.12'",
"fast-mlsirm==0.11.4",
]

[project.optional-dependencies]
Expand Down
25 changes: 20 additions & 5 deletions requirements.lock
Original file line number Diff line number Diff line change
Expand Up @@ -12,9 +12,9 @@ annotated-types==0.7.0 \
--hash=sha256:1f02e8b43a8fbbc3f3e0d4f0f4bfc8131bcb4eebe8849b8e5c773f3a1c582a53 \
--hash=sha256:aff07c09a53a08bc8cfccb9c85b05f1aa9a2a6f23728d790723543408344ce89
# via pydantic
anyio==4.14.1 \
--hash=sha256:4e5533c5b8ff0a24f5d7a176cbe6877129cd183893f66b537f8f227d10527d72 \
--hash=sha256:8d648a3544c1a700e3ff78615cd679e4c5c3f149904287e73687b2596963629e
anyio==4.14.2 \
--hash=sha256:9f505dda5ac9f0c8309b5e8bd445a8c2bf7246f3ce950121e45ea15bc41d1494 \
--hash=sha256:cfa139f3ed1a23ee8f88a145ddb5ac7605b8bbfd8592baacd7ce3d8bb4313c7f
# via
# httpx
# starlette
Expand Down Expand Up @@ -363,7 +363,20 @@ egressweave==0.1.0 \
--hash=sha256:6bcb07109bdee25a6d49e5516f4c99ecd172ffe8536455d2cac860c44a4492f6 \
--hash=sha256:e1e3a6dbabd4084fb03f19a95931ab96e4beeef96bc8fc7cf0d8e5b91e266057
# via contextual-orchestrator (pyproject.toml)
fast-mlsirm @ git+https://github.com/ContextualWisdomLab/fast-mlsirm.git@09f762ded35786dd1078222a4577ff09d649816f
fast-mlsirm==0.11.4 \
--hash=sha256:1779228af4c5d932514e465e57e137fd88f0d7e51d9463e69f1c6efa901bc179 \
--hash=sha256:1ce70211d8f837e489d456863cba5c163d0d27972ecc1689b7787af0b40daed6 \
--hash=sha256:1e65070ffab835a8af7e3eef3f67740b5146e9325d1a52b84b72f30e8b23998e \
--hash=sha256:24bf3f0edd8a23a53a3e6ff4e9291f4574644c6e10b44c43294703dc97f0ebf8 \
--hash=sha256:761f13783484ba7ef82c810c9c740f113d20cf52f879556056dfae9b36e39731 \
--hash=sha256:9ee76243c95a548a5f988c070588256d8e8d3bc747356991588a1f0a22b04a2c \
--hash=sha256:a42038ee72a23220b3153153b3e290f454f5dee098ee0fbfc17d94aefbb215fa \
--hash=sha256:c3aaa8a526de58842dd2cddc759e25b765da3b6a2c36257de66db35ef191e664 \
--hash=sha256:c7af04c0d01e1a8c802d39278fac6edff7263f67a039a6e293fcbe42ecb440a1 \
--hash=sha256:d026033bc4f649534c1f4d9516c7cd081fc064bd292a4e2d302acdb9f2ee060e \
--hash=sha256:f010df97a1f8bc7e4800c1b341f40204e8b42ac467a4d96948886144a3a3575a \
--hash=sha256:f0980d90d2094ef01d482915b1786b8654445fde5d5ba97446eedac573aca1e0 \
--hash=sha256:f4923dfae8db5407d20cc9ce0a08baaf6273824a22d01709b4b0cbc8ccfdbb87
# via contextual-orchestrator (pyproject.toml)
fastapi==0.138.2 \
--hash=sha256:6432359d067a432134620e7c5e4c6e5063e7f37815bbbbf20acef14b0d2e3fc8 \
Expand Down Expand Up @@ -453,7 +466,9 @@ greenlet==3.5.5 \
--hash=sha256:f2e3d061b8e13aec2f0441689b3c71b244a20e5d274a52cb0f7e31bd1d139552 \
--hash=sha256:f7278591501941bb2456af102bb9cd59aab48c6cfd6e2dd68fa1290bb0c49a42 \
--hash=sha256:fef01bd457f11fc158b130ca0027a3c365693280e8e231b65bdaf57999f39f5b
# via contextual-orchestrator (pyproject.toml)
# via
# contextual-orchestrator (pyproject.toml)
# sqlalchemy
h11==0.16.0 \
--hash=sha256:4e35b956cf45792e4caa5885e69fba00bdbc6ffafbfa020300e549b208ee5ff1 \
--hash=sha256:63cf8bbe7522de3bf65932fda1d9c2772064ffb3dae62d55932da54b31cb6c86
Expand Down
1 change: 1 addition & 0 deletions rust-toolchain.toml
Original file line number Diff line number Diff line change
@@ -1,3 +1,4 @@
[toolchain]
channel = "1.97.1"
profile = "minimal"
components = ["rustfmt", "clippy"]
6 changes: 3 additions & 3 deletions tests/test_fast_mlsirm_runtime_contract.py
Original file line number Diff line number Diff line change
Expand Up @@ -7,15 +7,15 @@


def test_supported_python_floor_matches_fast_mlsirm_runtime() -> None:
"""Every supported interpreter must install the mandatory psychometric runtime."""
"""Every supported interpreter must install the released psychometric runtime."""
project_data = tomllib.loads(Path("pyproject.toml").read_text(encoding="utf-8"))["project"]

assert project_data["requires-python"] == ">=3.12"
fast_mlsirm_dependencies = [
dependency
for dependency in project_data["dependencies"]
if dependency.startswith("fast-mlsirm ")
if dependency.startswith("fast-mlsirm")
]
assert fast_mlsirm_dependencies == [
"fast-mlsirm @ git+https://github.com/ContextualWisdomLab/fast-mlsirm.git@09f762ded35786dd1078222a4577ff09d649816f ; python_full_version >= '3.12'"
"fast-mlsirm==0.11.4"
]
20 changes: 20 additions & 0 deletions tests/test_multi_model_policy_authority.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,20 @@
"""Fail-closed boundary for unreleased multi-model decision policy."""

from __future__ import annotations

import importlib

import pytest


@pytest.mark.parametrize(
"module_name",
(
"contextual_orchestrator.domain.combination",
"contextual_orchestrator.application.fan_out",
),
)
def test_uncalibrated_combination_policy_is_not_importable(module_name: str) -> None:
"""Do not ship selection or allocation without released owner evidence."""
with pytest.raises(ModuleNotFoundError):
importlib.import_module(module_name)
21 changes: 18 additions & 3 deletions uv.lock

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

Loading