Skip to content
Closed
Show file tree
Hide file tree
Changes from 250 commits
Commits
Show all changes
281 commits
Select commit Hold shift + click to select a range
1c6819d
docs(adr): require terminal merge checks
Aug 11, 2026
06e6ffb
docs(adr): record Strix fail-closed gate
Aug 11, 2026
338c11d
test(adr): require enforced exact-head merge controls
seonghobae Aug 11, 2026
03c6630
docs(adr): enforce exact-head merge controls
seonghobae Aug 11, 2026
6a7e38e
docs(adr): record fast numerical stability fix
Aug 11, 2026
70fc9d8
docs(adr): reconcile fast stability fix with main
Aug 11, 2026
574fa8c
docs(adr): refresh exact-head CI evidence
Aug 11, 2026
a42b7b6
docs(adr): record current-head check binding
Aug 11, 2026
77c2186
docs: record CI portability remediation
Aug 11, 2026
0ad4193
docs: refresh exact-head CI snapshot
Aug 11, 2026
49002cc
docs: record dependency floor guard
Aug 11, 2026
9b66a13
docs: record larger local judge reliability sweep
Aug 11, 2026
806bda5
docs: record external GitHub API incident context
Aug 11, 2026
c3027c1
docs(adr): record strict fast judge schema hardening
Aug 11, 2026
2aeab27
docs(adr): harden review evidence publication
Aug 11, 2026
0f37a76
docs(adr): record central gate remediation
Aug 11, 2026
5163030
docs(adr): record non-compliant merge incident
Aug 11, 2026
b67c051
docs(adr): record central review follow-up gaps
Aug 11, 2026
0263ad5
docs(adr): keep dependabot discrepancy visible
Aug 11, 2026
afe0956
docs: record OA PDF attachment revalidation
Aug 11, 2026
3d42c8d
docs: record archived OA PDF retrieval evidence
Aug 11, 2026
877c77f
docs: record draft review blocker
Aug 11, 2026
8d629f4
fix: make ledger SQL templates scanner-clean
Aug 11, 2026
895eb22
test: enforce exact-head ADR controls
Aug 11, 2026
4bd9c2b
docs: record exact-head evidence correction
Aug 11, 2026
1b7dbd2
docs: record cumulative threshold calibration
Aug 11, 2026
3781930
docs: record local judge framing reliability
seonghobae Aug 11, 2026
b65711f
docs: expand merge gate evidence requirements
seonghobae Aug 12, 2026
6b6edba
docs: record cross-repo Strix false-green evidence
seonghobae Aug 12, 2026
c09cac0
docs: record central Strix provenance gap
seonghobae Aug 12, 2026
2a4270a
docs: separate routing hints from model judgment
seonghobae Aug 12, 2026
125bfc1
docs: record central review cleanup findings
seonghobae Aug 12, 2026
32ba3a9
docs: record fast Strix evidence boundary
seonghobae Aug 12, 2026
c138d17
docs: close central cleanup remediation record
seonghobae Aug 12, 2026
ddd0591
docs: record local judge retry calibration
seonghobae Aug 12, 2026
bbb6361
feat: bridge Codex Responses to local MLX
seonghobae Aug 12, 2026
ab0c05c
docs: record workflow contract reachability finding
seonghobae Aug 12, 2026
fccac33
docs: record live gateway judge method sensitivity
seonghobae Aug 12, 2026
54dfce2
feat: expose orchestrator control-plane candidate
seonghobae Aug 12, 2026
21a17f0
fix: separate candidate discovery from disablement
seonghobae Aug 12, 2026
13fe555
feat: expose full local candidate registry
seonghobae Aug 12, 2026
a107892
fix(local): avoid retry multiplication on mlx queues
seonghobae Aug 12, 2026
8394373
fix(api): reject unknown passthrough models
seonghobae Aug 12, 2026
da24b1a
fix(api): separate model readiness from registry status
seonghobae Aug 12, 2026
583c4f3
Merge remote-tracking branch 'origin/codex/local-llm-benchmark' into …
seonghobae Aug 12, 2026
c1c7997
docs(adr): record network and ruleset merge boundaries
seonghobae Aug 12, 2026
216177f
fix(local): honor provider-specific retry budgets
seonghobae Aug 12, 2026
56d3085
docs(adr): correct draft review-gate status
seonghobae Aug 12, 2026
8fee52b
fix(local): reject malformed responses and retry config
seonghobae Aug 12, 2026
5b49594
docs(adr): refresh exact-head coverage evidence
seonghobae Aug 12, 2026
1ffdc3f
docs(adr): bind review note to latest head
seonghobae Aug 12, 2026
ea254c2
docs(adr): clarify implementation head provenance
seonghobae Aug 12, 2026
2e76caa
feat(mlsirm): route local judge to contextual-orchestrator and add po…
seonghobae Aug 13, 2026
ea4dfa9
fix: enforce fast judge and multi-item evidence
seonghobae Aug 13, 2026
46725e2
fix: bind provider allowlists to explicit config
seonghobae Aug 13, 2026
c7b5dbc
docs: record mlx model and batch throughput probe
seonghobae Aug 13, 2026
79ce905
docs(adr): record current Strix false-green evidence
seonghobae Aug 13, 2026
4af8e26
docs(adr): record trusted Strix evidence follow-ups
seonghobae Aug 13, 2026
f7b2079
docs(adr): distinguish VPN LibreSSL path failures
seonghobae Aug 13, 2026
17d636b
docs(adr): record strix smoke-contract regression
seonghobae Aug 13, 2026
b7ee4bb
docs(adr): record hyphenated strix marker guard
seonghobae Aug 13, 2026
becc758
docs(benchmark): record repeated mlx concurrency probe
seonghobae Aug 13, 2026
a1d8128
fix(local): bound batch concurrency input
seonghobae Aug 13, 2026
fe7b632
docs(local): document concurrency bound
seonghobae Aug 13, 2026
8a88063
docs(judge): record current mlx K sweep
seonghobae Aug 13, 2026
5ea0f77
docs(adr): record runner-capacity merge blocker
seonghobae Aug 13, 2026
d3ed441
docs(adr): record unbound Strix status fallback
seonghobae Aug 13, 2026
a159dff
docs(adr): record fresh runner-capacity blocker
seonghobae Aug 13, 2026
4d8cd86
fix(judge): require fast-mlsirm calibration boundary
seonghobae Aug 13, 2026
998c86f
test(judge): isolate missing calibration dependency
seonghobae Aug 13, 2026
fb16482
fix(ci): quote fuzz time budget expansion
seonghobae Aug 13, 2026
76de3e1
docs(adr): record unbound strix metadata finding
seonghobae Aug 13, 2026
51ac452
docs(adr): record independent review authority boundary
seonghobae Aug 13, 2026
ec94d99
docs(adr): record false-green Strix artifact
seonghobae Aug 13, 2026
7cb34a9
docs(adr): record current Strix false-green evidence
seonghobae Aug 13, 2026
c646ba5
docs(bench): record 3b mlx saturation probe
seonghobae Aug 13, 2026
5d80e11
perf(local): reuse bounded concurrency for in-process batches
seonghobae Aug 13, 2026
454b9f5
docs(adr): qualify local batch error propagation
seonghobae Aug 13, 2026
b81f8d7
test(batch): assert concurrent result order
seonghobae Aug 13, 2026
173288c
docs(bench): record coordinator mlx smoke
seonghobae Aug 13, 2026
f7141ea
docs(adr): record latest CI capacity evidence
seonghobae Aug 13, 2026
361725b
fix(runtime): guard circuit breaker under batch concurrency
seonghobae Aug 13, 2026
d3abf51
docs(bench): record warm coordinator smoke
seonghobae Aug 13, 2026
8aaa0e0
docs(adr): record latest central Strix false green
seonghobae Aug 13, 2026
823ce91
docs(adr): record fast Strix base evidence
seonghobae Aug 13, 2026
c5e0286
docs(adr): record CI cancellation evidence
seonghobae Aug 13, 2026
4b954e4
docs(adr): record approval policy drift
seonghobae Aug 13, 2026
73c6c98
docs(adr): record semantic judge miss
seonghobae Aug 13, 2026
e59ef71
docs(adr): record latest ci capacity evidence
seonghobae Aug 13, 2026
a29ff9e
docs(calibration): record paired binary threshold evidence
seonghobae Aug 13, 2026
1eef1e9
docs(calibration): record bounded binary concurrency
seonghobae Aug 13, 2026
9bc620d
perf(mlx): align HTTP admission with measured batches
seonghobae Aug 13, 2026
0f61128
fix(server): bound concurrent run configuration
seonghobae Aug 13, 2026
78894f6
docs(adr): record exact-head CI cancellations
seonghobae Aug 13, 2026
d82e592
perf(judge): expose gateway concurrency capability
seonghobae Aug 13, 2026
9b6dec3
docs(benchmark): record integrated judge smoke
seonghobae Aug 13, 2026
d8c1b73
docs(adr): record polytomous default hardening
seonghobae Aug 13, 2026
a0a354a
docs(adr): record current PR gate state
seonghobae Aug 13, 2026
618810e
docs(benchmark): record integrated polytomous smoke
seonghobae Aug 13, 2026
db45804
docs(adr): retain integrated judge failure evidence
seonghobae Aug 13, 2026
f8cac36
docs(benchmark): record anchored model and throughput evidence
seonghobae Aug 13, 2026
6422a20
docs(adr): record anchored judge merge gate
seonghobae Aug 13, 2026
18d8c3b
docs(adr): record synchronized judge head
seonghobae Aug 13, 2026
e0000ad
docs(benchmark): record MLX readiness and judge failure
seonghobae Aug 13, 2026
c518049
docs(adr): record latest exact-head gate
seonghobae Aug 13, 2026
c2bb2b2
feat: add bounded provider readiness refresh
seonghobae Aug 13, 2026
2f904a2
docs: record anchored judge calibration reruns
seonghobae Aug 13, 2026
bd4c1a3
docs: reconcile linked fast judge head
seonghobae Aug 13, 2026
fe8437d
docs: track fast judge evidence head
seonghobae Aug 13, 2026
bcb55d0
docs: record exact-head comment correction
seonghobae Aug 13, 2026
c860644
docs: record protection policy reconciliation
seonghobae Aug 13, 2026
273943e
docs: distinguish hosted runner queue evidence
seonghobae Aug 13, 2026
b782b01
docs: record repeated mlx concurrency sweep
seonghobae Aug 13, 2026
a7de9f6
docs: record integrated two-item judge smoke
seonghobae Aug 13, 2026
96501bb
docs: record latest exact-head merge audit
seonghobae Aug 13, 2026
910d9ef
docs: record post-main mlx smoke
seonghobae Aug 13, 2026
9d4562f
docs: record option-only bias literature
seonghobae Aug 13, 2026
435c3ef
docs: bind corrected exact-head review evidence
seonghobae Aug 13, 2026
153ca6b
docs: record full-sha review evidence correction
seonghobae Aug 13, 2026
fa20fe3
docs: record paired option-count calibration
seonghobae Aug 13, 2026
2f42a1b
docs: extend model calibration and merge evidence
seonghobae Aug 13, 2026
d83a029
docs: bind benchmark to final judge head
seonghobae Aug 13, 2026
d1f97a2
docs: record LibreSSL transport diagnosis
seonghobae Aug 13, 2026
c460270
docs: bind linked judge head to merge audit
seonghobae Aug 13, 2026
36b3d2a
docs: record latest linked judge head
seonghobae Aug 13, 2026
cc0d295
docs: record current CI capacity gate
seonghobae Aug 13, 2026
54ef9bb
docs: record binary judge calibration evidence
seonghobae Aug 13, 2026
237ba2d
docs: record judge concurrency plateau
seonghobae Aug 13, 2026
01aafd5
docs: reconcile OA and hosted gate evidence
seonghobae Aug 13, 2026
f783e95
docs: record effective protection drift
seonghobae Aug 13, 2026
bb88bf7
docs: record protection remediation
seonghobae Aug 13, 2026
7f47665
fix: require gateway provenance for LLM judges
seonghobae Aug 13, 2026
c722fa8
docs: record judge saturation evidence
seonghobae Aug 13, 2026
f3d54b1
docs: record occupancy merge gate
seonghobae Aug 13, 2026
af500ee
docs: record fast base synchronization
seonghobae Aug 13, 2026
cb88960
docs: record hosted runner capacity gate
seonghobae Aug 13, 2026
1614c7f
docs: record draft review gate
seonghobae Aug 13, 2026
adc4f80
chore: keep unreliable tiny model out of verifier lane
seonghobae Aug 13, 2026
67ed128
docs: record verifier routing gate
seonghobae Aug 13, 2026
e5180c3
docs: pin latest verifier gate evidence
seonghobae Aug 13, 2026
490dfd8
fix: route unreliable mlx judges away from verifier
seonghobae Aug 13, 2026
0df4267
docs: record final verifier routing merge gate
seonghobae Aug 14, 2026
cc413a4
docs: record protection drift remediation
seonghobae Aug 14, 2026
f6456e1
docs: record final exact-head audit
seonghobae Aug 14, 2026
8ecd379
fix: bound local MLX model switching
seonghobae Aug 14, 2026
3f33a72
feat: add fast-mlsirm runtime preflight
seonghobae Aug 14, 2026
2816140
docs: record OA PDF rights conflict
seonghobae Aug 14, 2026
d3480cc
docs: gate divergent local-llm PR behind security base
seonghobae Aug 14, 2026
6373d98
docs: record held-out MLX calibration limits
seonghobae Aug 14, 2026
63451a0
docs: refresh protected merge evidence
seonghobae Aug 14, 2026
6354c04
docs: record dedicated MLX calibration evidence
seonghobae Aug 14, 2026
62100d3
fix(local): validate MLX model registry before readiness probe
seonghobae Aug 14, 2026
bd2515f
docs: record 3B non-ceiling calibration result
seonghobae Aug 14, 2026
b36878c
docs: bind latest exact-head merge evidence
seonghobae Aug 14, 2026
52e0746
docs: reject unbound Strix evidence
seonghobae Aug 14, 2026
29e1388
docs: bind central Strix dependency
seonghobae Aug 14, 2026
b30697d
docs(adr): record stale gates and calibration stratification
seonghobae Aug 14, 2026
160471b
docs(calibration): record K-stratified MLX evidence
seonghobae Aug 14, 2026
a5f7697
docs: record unbound Strix evidence
seonghobae Aug 14, 2026
7cfad1b
docs: record review evidence transport fix
seonghobae Aug 14, 2026
d00431e
docs: record Strix provider recovery
seonghobae Aug 14, 2026
4465922
docs: record organization ruleset governance gap
seonghobae Aug 14, 2026
da849d3
docs: record repaired organization policy
seonghobae Aug 14, 2026
27aa4ad
docs: record review identity correction
seonghobae Aug 14, 2026
b2b3d8e
docs: record provider execution contract failure
seonghobae Aug 14, 2026
d1bc819
docs: clarify dedicated MLX benchmark endpoint
seonghobae Aug 14, 2026
8bf757f
docs: record exact-head identity correction
seonghobae Aug 14, 2026
8194e04
docs: record current unbound security evidence
seonghobae Aug 14, 2026
383ccfd
docs: record contextual strix provenance gap
seonghobae Aug 14, 2026
50b91c3
docs: clarify intermittent libreSSL transport evidence
seonghobae Aug 14, 2026
792c9ce
docs: record central evidence boundary remediation
seonghobae Aug 14, 2026
b72a838
docs: record exact-head comment correction
seonghobae Aug 14, 2026
e99b097
docs: record structured status regressions
seonghobae Aug 14, 2026
abfb45a
docs: record unbound central strix success
seonghobae Aug 14, 2026
9f45226
docs: record stale review request correction
seonghobae Aug 14, 2026
eaf7ff3
docs: record safe review payload transport
seonghobae Aug 14, 2026
bad1e1a
fix: prevent concurrent completion ID collisions
seonghobae Aug 14, 2026
a07c11f
docs: record review transport limits
seonghobae Aug 14, 2026
19c3e88
docs: record fast judge contract preflight regression
seonghobae Aug 14, 2026
4923447
docs: record linked PR runner gate state
seonghobae Aug 14, 2026
a9278d1
docs: record judge evidence redaction
seonghobae Aug 14, 2026
c0c2ecb
docs: bind current judge preflight evidence
seonghobae Aug 14, 2026
73853a4
docs: bind final linked PR gate evidence
seonghobae Aug 14, 2026
3ab44dd
docs: correct runner evidence interpretation
seonghobae Aug 14, 2026
474b667
docs: record warm mlx gateway throughput
seonghobae Aug 14, 2026
1d3e062
docs: record exact-head judge smoke
seonghobae Aug 14, 2026
1b22ff0
fix: retry transient TLS socket failures
seonghobae Aug 14, 2026
bc882c0
docs: pin TLS retry ADR evidence
seonghobae Aug 14, 2026
0fb8fdb
docs: record exact-head K calibration
seonghobae Aug 14, 2026
f0f30f3
docs: record current actions queue evidence
seonghobae Aug 14, 2026
3bc1264
docs: record review identity correction
seonghobae Aug 14, 2026
8f922d8
docs: distinguish actions queue from outage
seonghobae Aug 14, 2026
1c5b5c8
docs: record current MLX judge recheck
seonghobae Aug 14, 2026
9f662c4
docs: clarify cross-repository judge runtime
seonghobae Aug 14, 2026
070d929
docs: record direct MLX gateway comparison
seonghobae Aug 14, 2026
53f47a6
docs: record current integrated judge smoke
seonghobae Aug 14, 2026
dac931f
docs: record unbound exact-head strix evidence
seonghobae Aug 14, 2026
229bb30
fix: redact provider probe diagnostics
seonghobae Aug 14, 2026
f15ccb0
docs: record warm local mlx gateway recheck
seonghobae Aug 14, 2026
c6f30eb
docs: record exact-head judge smoke
seonghobae Aug 14, 2026
228126b
docs: align container secret guidance with KV
seonghobae Aug 14, 2026
cdca9d8
fix: reject incomplete batch responses
seonghobae Aug 14, 2026
018f6ef
docs: record current-head judge smoke
seonghobae Aug 14, 2026
c580839
docs: record current hosted gate state
seonghobae Aug 14, 2026
61f94a0
docs: record unbound strix artifact
seonghobae Aug 14, 2026
c179b54
docs: record exact Strix artifact identity gate
seonghobae Aug 14, 2026
ccfa292
docs: record exact-head comment correction
seonghobae Aug 14, 2026
15a16d8
docs: record current MLX integrated smoke
seonghobae Aug 14, 2026
06dfa40
docs: record minimal Strix successor gate
seonghobae Aug 14, 2026
b1dc49e
docs: record inconsistent PR snapshot guard
seonghobae Aug 14, 2026
87e1c3b
docs: record central suite validation
seonghobae Aug 14, 2026
df5f5bf
docs: record quoted check audit paths
seonghobae Aug 14, 2026
a9cd011
docs: harden shell audit argument handling
seonghobae Aug 14, 2026
71164ae
docs: record current Actions queue evidence
seonghobae Aug 14, 2026
a139ece
docs: record superseded central PR closure
seonghobae Aug 14, 2026
1c05e20
docs: record current gateway throughput sweep
seonghobae Aug 14, 2026
bd78ca1
docs: record strict option-count calibration
seonghobae Aug 14, 2026
ebc76d4
docs: record current exact-head queue
seonghobae Aug 14, 2026
719d9cc
docs: compare polytomous judge methods
seonghobae Aug 14, 2026
e9935d7
docs: record routed mlx judge recheck
seonghobae Aug 14, 2026
aeae379
docs: record fresh local gateway evidence
seonghobae Aug 14, 2026
efc4b5f
docs: record current protected gate state
seonghobae Aug 14, 2026
1003cea
docs: record check materialization re-poll
seonghobae Aug 14, 2026
40b7bac
fix(security): enforce scoped token modes
seonghobae Aug 14, 2026
2e36f9b
fix(security): keep scope precedence at authorization
seonghobae Aug 14, 2026
491566e
docs: record vpn and trusted strix failure evidence
seonghobae Aug 14, 2026
887cc66
docs: record current protected merge gate
seonghobae Aug 14, 2026
cdfc15d
docs: record gateway rate-limit benchmark boundary
seonghobae Aug 14, 2026
fcb82a1
docs: record check audit quoting regression
seonghobae Aug 14, 2026
3a59b67
docs: record invalid review identity
seonghobae Aug 14, 2026
007acf6
docs: record current hosted gate state
seonghobae Aug 14, 2026
0b19112
docs: correct mlx benchmark evidence
seonghobae Aug 14, 2026
60d9cfc
fix: separate local gateway credentials and structured judge transport
seonghobae Aug 14, 2026
9058a08
docs: record hosted security provenance gaps
seonghobae Aug 14, 2026
e3e9788
docs: record Strix provider contract failure
seonghobae Aug 14, 2026
621c42f
docs: record MLX output budget calibration
seonghobae Aug 14, 2026
d7133c4
docs: record linked security finding
seonghobae Aug 14, 2026
64b6d56
docs: record linked fix validation
seonghobae Aug 14, 2026
e0413fe
docs: record unbound strix evidence
seonghobae Aug 14, 2026
68002d0
docs: record current Strix review evidence drift
seonghobae Aug 14, 2026
d5236bb
docs: bind latest contextual Strix evidence state
seonghobae Aug 14, 2026
de4d95a
docs: record Strix failure-marker evidence
seonghobae Aug 14, 2026
2acb3a4
docs: bind latest contextual zero-finding evidence
seonghobae Aug 14, 2026
775ea13
docs: record live orchestration latency
seonghobae Aug 14, 2026
badcf28
fix(ci): align atheris with coverage image
seonghobae Aug 15, 2026
6a411be
docs(adr): bind coverage review evidence
seonghobae Aug 15, 2026
d2072fa
docs(adr): record review dependency gate
seonghobae Aug 15, 2026
63d9abf
docs(adr): record unbound green Strix evidence
seonghobae Aug 15, 2026
83b17bd
docs: record unbound Strix SQL finding false positive
seonghobae Aug 15, 2026
a1d486a
docs: record current bound-free Strix result
seonghobae Aug 15, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 6 additions & 0 deletions .adr-config.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,6 @@
project_slug: contextual-orchestrator
owner: ContextualWisdomLab
default_status: proposed
decision_id_format: NNNN
template_source: madr-v4
last_decision_id: 0009
4 changes: 4 additions & 0 deletions .github/dependabot.yml
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,8 @@ updates:
day: monday
time: "04:00"
timezone: Asia/Seoul
cooldown:
default-days: 7
open-pull-requests-limit: 5

- package-ecosystem: pip
Expand All @@ -16,4 +18,6 @@ updates:
day: monday
time: "04:30"
timezone: Asia/Seoul
cooldown:
default-days: 7
open-pull-requests-limit: 5
11 changes: 7 additions & 4 deletions .github/workflows/fuzz.yml
Original file line number Diff line number Diff line change
Expand Up @@ -72,16 +72,19 @@ jobs:
fi

- name: Fuzz request-body parser
run: python fuzz/fuzz_request_body.py -max_total_time=${FUZZ_SECONDS} -artifact_prefix=crash- fuzz/corpus/request_body
run: python fuzz/fuzz_request_body.py -max_total_time="${FUZZ_SECONDS}" -artifact_prefix=crash- fuzz/corpus/request_body

- name: Fuzz agent-config parser
run: python fuzz/fuzz_agent_config.py -max_total_time=${FUZZ_SECONDS} -artifact_prefix=crash- fuzz/corpus/agent_config
run: python fuzz/fuzz_agent_config.py -max_total_time="${FUZZ_SECONDS}" -artifact_prefix=crash- fuzz/corpus/agent_config

- name: Fuzz secret redaction
run: python fuzz/fuzz_redaction.py -max_total_time=${FUZZ_SECONDS} -artifact_prefix=crash- fuzz/corpus/redaction
run: python fuzz/fuzz_redaction.py -max_total_time="${FUZZ_SECONDS}" -artifact_prefix=crash- fuzz/corpus/redaction

- name: Fuzz orchestration engine
run: python fuzz/fuzz_orchestration.py -max_total_time=${FUZZ_SECONDS} -artifact_prefix=crash- fuzz/corpus/orchestration
run: python fuzz/fuzz_orchestration.py -max_total_time="${FUZZ_SECONDS}" -artifact_prefix=crash- fuzz/corpus/orchestration

- name: Fuzz model-judge response parser
run: python fuzz/fuzz_model_judge.py -max_total_time="${FUZZ_SECONDS}" -artifact_prefix=crash- fuzz/corpus/judge

- name: Upload crash artifacts
if: failure()
Expand Down
5 changes: 3 additions & 2 deletions CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -44,8 +44,9 @@ python fuzz/fuzz_request_body.py -max_total_time=60 fuzz/corpus/request_body
python -m contextual_orchestrator "your prompt" --agents examples/agents.mock.json

# Serve the OpenAI-compatible API + /admin console
export CONTEXTUAL_ORCHESTRATOR_TOKEN="$(python -c 'import secrets; print(secrets.token_urlsafe(32))')"
python -m contextual_orchestrator --serve --agents examples/agents.mock.json --port 8000
local_token="$(python -c 'import secrets; print(secrets.token_urlsafe(32))')"
python -m contextual_orchestrator --serve --agents examples/agents.mock.json --port 8000 \
--auth-token "$local_token"

# Loopback-only local dev server (auth disabled; loopback only)
./.superset/run.sh
Expand Down
11 changes: 6 additions & 5 deletions Dockerfile
Original file line number Diff line number Diff line change
Expand Up @@ -2,10 +2,11 @@
# tree on a slim Python base. Runs the OpenAI-compatible server.
#
# Build: docker build -t contextual-orchestrator .
# Run : docker run --rm -p 8000:8000 \
# -e CONTEXTUAL_ORCHESTRATOR_TOKEN=change-me \
# -e OPENAI_API_KEY=sk-... \
# contextual-orchestrator
# Run : seed CONTEXTUAL_ORCHESTRATOR_TOKEN and provider credentials into the KV
# registry first, then use:
# docker run --rm -p 8000:8000 contextual-orchestrator
# Runtime secrets are never passed through the container environment or argv;
# see docs/kv-credentials.md for the bootstrap flow.
# Agents: defaults to the bundled mock pool; mount your own and set AGENTS_FILE:
# -v ./agents.json:/app/agents.json -e AGENTS_FILE=/app/agents.json
# python:3.12-slim
Expand All @@ -26,4 +27,4 @@ HEALTHCHECK --interval=30s --timeout=3s --start-period=5s \
CMD ["python", "-c", "import urllib.request,os;urllib.request.urlopen(f'http://127.0.0.1:{os.environ.get(\"PORT\",\"8000\")}/healthz', timeout=2)"]

# --allow-public-bind: 컨테이너 내부 0.0.0.0 바인딩 필요(외부 노출은 호스트 포트 매핑이 결정)
CMD ["sh", "-c", "python -m contextual_orchestrator --serve --agents \"$AGENTS_FILE\" --host 0.0.0.0 --port \"$PORT\" --allow-public-bind"]
CMD ["sh", "-c", "python -m contextual_orchestrator --serve --agents \"$AGENTS_FILE\" --host 0.0.0.0 --port \"$PORT\" --allow-public-bind --auth-token-key CONTEXTUAL_ORCHESTRATOR_TOKEN"]
62 changes: 54 additions & 8 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,7 +5,11 @@

Stdlib Python lab for a single API that routes, delegates, verifies, and synthesizes work across a configurable pool of OpenAI-compatible model agents.

This is not a Sakana AI product or a reproduction of their trained models. It is a small implementation of the public architecture pattern: expose one model-like interface while keeping the agent pool, routing, workflow, and verification logic behind it.
This is not a Sakana AI product or a reproduction of their trained models. It is
a small implementation of the public architecture pattern: expose one
model-like orchestration candidate while keeping the worker pool, routing,
workflow, and verification logic behind it. `contextual-orchestrator` is the
public control-plane model; it is not just an HTTP gateway.

## Quick Start

Expand All @@ -17,8 +21,9 @@ python -m contextual_orchestrator "Summarize why model orchestration helps long
Run the OpenAI-compatible subset:

```bash
export CONTEXTUAL_ORCHESTRATOR_TOKEN="$(python -c 'import secrets; print(secrets.token_urlsafe(32))')"
python -m contextual_orchestrator --serve --agents examples/agents.mock.json --port 8000
local_token="$(python -c 'import secrets; print(secrets.token_urlsafe(32))')"
python -m contextual_orchestrator --serve --agents examples/agents.mock.json --port 8000 \
--auth-token "$local_token"
```

Admin console:
Expand All @@ -29,14 +34,15 @@ http://127.0.0.1:8000/admin

```bash
curl -s http://127.0.0.1:8000/v1/chat/completions \
-H "authorization: Bearer $CONTEXTUAL_ORCHESTRATOR_TOKEN" \
-H "authorization: Bearer $local_token" \
-H 'content-type: application/json' \
-d '{"model":"contextual-orchestrator","messages":[{"role":"user","content":"Analyze this code review task and verify the answer."}]}' | jq .
```

HTTP serving is hardened for local lab use:

- `/admin`, `/admin/state`, `/api/v1/*`, and `/v1/chat/completions` require a Bearer token. Use `--admin-token` and `--inference-token` to separate operator and runtime access, or `--auth-token` / `CONTEXTUAL_ORCHESTRATOR_TOKEN` for one local-development token.
- `/admin`, `/admin/state`, `/api/v1/*`, and `/v1/chat/completions` require a Bearer token. Use `--admin-token-key` and `--inference-token-key` to resolve split tokens from the KV, or `--auth-token-key` for one token. Explicit `--auth-token`/split-token values are local-development escape hatches; the CLI no longer reads auth secrets from environment variables.
- A production deployment that uses the ecosystem identity plane must inject a reviewed `bearer_verifier` into `SecurityConfig` to validate Keyverse-issued OIDC tokens (issuer, audience, signature, expiry, and scope). The core does not hand-roll JWT parsing or hold Keycloak admin credentials; a static bearer token is not a Keyverse integration.
- Binding to `0.0.0.0` or `::` requires `--allow-public-bind`.
- JSON request bodies, chat message roles, orchestration modes, body sizes, request rate, and concurrent run counts are validated before orchestration runs.
- Full orchestration traces are not returned by default. Set `include_orchestration_trace: true` per chat request or start with `--expose-trace-by-default` when the caller is trusted.
Expand All @@ -60,6 +66,36 @@ Use real workers by replacing `mock://` agents with OpenAI-compatible endpoints.
}
```

For a local `mlx-lm` OpenAI-compatible server, use the explicit `mlx://` scheme. It is loopback-only, does not require a credential, and is translated to HTTP only after the loopback check:

```json
{
"agents": [
{
"id": "local_fast_agent",
"model": "mlx-community/llama-3.2-3b-instruct-4bit",
"base_url": "mlx://127.0.0.1:8080/v1",
"provider_name": "mlx-lm",
"tags": ["reasoning", "coding", "verification"]
}
]
}
```

The full local candidate registry is [examples/agents.local.json](examples/agents.local.json).
It contains the public `contextual-orchestrator` candidate, discovered MLX
worker models, and every discovered llama.cpp/LM Studio candidate. Discovery
does not decide governance state: seed candidates are enabled by default, while
`disabled` is reserved for an explicit operator/admin quarantine or a persisted
removal tombstone. The contextual-orchestrator record is excluded from internal
roles because this implementation has no bounded recursive self-call protocol;
that is a routing safety constraint, not a disabled candidate. The registry is
explicit; runtime discovery does not silently change the pool.

Run an evaluation against that server with `--temperature 0` for repeatable judging. For reasoning-capable mlx models, pass `--chat-template-args '{"enable_thinking":false}'` when a short structured judge response is required. `--local-concurrency N` enables bounded concurrent local batch requests (`1..64`; the current measured starting point for this server is `8`); when serving HTTP, set `--max-concurrent-runs N` explicitly as well if the measured batch concurrency exceeds the secure default of `8`. Keep interactive route/conduct requests on the default sequential path.

Model-based conduct verification requires `fast-mlsirm` in the same runtime and fails closed when it is absent or broken; fast-mlsirm sends its judge completion through this contextual-orchestrator gateway, so no direct provider fallback is used. “Same runtime” means that the exact interpreter used for the live run can import both packages: install both checkouts into one environment (prefer editable installs), or expose both source roots with `PYTHONPATH` during a source run. Before a live judge benchmark, run `python -m contextual_orchestrator check-fast-mlsirm` with that exact interpreter. It prints the interpreter, package version, transitive-import status, and contextual contract check, and exits nonzero on a missing dependency or contract mismatch. Do not run the preflight in one virtual environment and the judge in another. See [ADR 0001](docs/planning/adrs/0001-fail-closed-model-judgment.md).

The agent pool is manageable at runtime: `POST`/`PATCH`/`DELETE` on `/api/v1/agent_pools/default/worker_agents[/{id}]` add, govern, and remove model-group members. Pass `--agents-db PATH` (or `CONTEXTUAL_ORCHESTRATOR_AGENTS_DB`) to persist those changes to a stdlib sqlite file — stored changes overlay the seed agents file at startup, and removals write disabled tombstones so they survive restarts; without it the pool is in-memory as before.

Seed the credential into the KV once at bootstrap:
Expand All @@ -68,14 +104,21 @@ Seed the credential into the KV once at bootstrap:
echo "$OPENAI_API_KEY" | python -m contextual_orchestrator register-credential --name OPENAI_API_KEY --value-stdin
```

Non-mock providers must use `https://` URLs and a **resolvable KV credential** — a non-mock agent whose credential is missing raises `NotConfigured` rather than falling back to an environment variable. The runtime blocks loopback, private, link-local, multicast, and reserved provider addresses before sending a key. Set `CONTEXTUAL_ORCHESTRATOR_ALLOWED_PROVIDER_HOSTS` to a comma-separated host allowlist when only approved model gateways should be reachable. External calls use a timeout and default output token cap.
For a persistent KV-backed server token, seed a credential such as
`CONTEXTUAL_ORCHESTRATOR_TOKEN` and start with
`--auth-token-key CONTEXTUAL_ORCHESTRATOR_TOKEN`. The in-memory credential
backend is process-local and is suitable only for tests; production auth
registration and OIDC client secrets belong to the deployment/KV boundary.

Non-mock providers must use `https://` URLs and a **resolvable KV credential** — a non-mock agent whose credential is missing raises `NotConfigured` rather than falling back to an environment variable. The runtime blocks loopback, private, link-local, multicast, and reserved provider addresses before sending a key. Pass `--allowed-provider-host HOST` once per approved gateway when an explicit host allowlist is required; it is bound at client construction and is not changed by request-time environment variables. External calls use a timeout and default output token cap.

> The legacy `api_key_env` field is still accepted for back-compat, but its value is now treated as the **credential name** in the KV, not as an environment variable to read. This supersedes the old `api_key_env` env pattern.

## Architecture

One public interface:

- `contextual-orchestrator` is the model-like control-plane candidate exposed to callers. `/v1/models` lists it first, followed by every configured worker candidate, including disabled candidates with their status.
- `/v1/chat/completions` accepts normal chat messages, and `"stream": true` returns an OpenAI-compatible `text/event-stream` of `chat.completion.chunk` deltas terminated by `data: [DONE]`. In **route** mode the worker's tokens are streamed live as they arrive from the provider (real token streaming); in **conduct** mode the multi-step answer is produced then framed as deltas (a workflow can't honestly token-stream a synthesizer that hasn't run yet).
- `TaskOrchestrator.complete()` decides whether to route to one worker or run a short workflow.
- `TaskOrchestrator.compare_to_baseline(prompts, mode)` (CLI `--eval PROMPT...`) measures the orchestration engine against a single-worker baseline — per-prompt and aggregate latency plus a structural coverage delta (contributing steps + verifier-pass presence). It is a measured tradeoff report, not a human-quality claim.
Expand Down Expand Up @@ -125,7 +168,7 @@ Local spend observability, aggregated from in-memory workflow runs. It is honest

```bash
curl -s http://127.0.0.1:8000/api/v1/spend_analytics/latest \
-H "authorization: Bearer $CONTEXTUAL_ORCHESTRATOR_TOKEN" | jq '.totals, .by_model, .budget'
-H "authorization: Bearer $local_token" | jq '.totals, .by_model, .budget'
```

- **Tokens.** `by_model[].output_tokens` uses the provider-reported `usage.completion_tokens` when a real worker returns it, and falls back to a `~4 chars/token` estimate otherwise. Each row carries `usage_source`: `reported` (all steps reported), `mixed`, or `estimated`. `estimated_output_tokens` is always the estimate, kept alongside for comparison. `measurement_status` is `local_runtime_estimate`, not production telemetry.
Expand Down Expand Up @@ -192,7 +235,10 @@ is read from a **KV config store**, never `os.getenv`.
backend (local in-process backend standalone), and records one usage-ledger row
per original vector with the full attribution dimensions (service, team,
group, company, provider) carried in `metadata`.
- **Health.** `GET /healthz` is an unauthenticated liveness probe.
- **Health.** `GET /healthz` is an unauthenticated liveness probe; it never
claims that an upstream chat worker is serving. Admins can use
`GET /api/v1/provider_readiness/latest?refresh=true` for one bounded,
non-retrying chat probe per enabled worker.
- **Standalone + optional pg-llm-batch integration.** The hub runs standalone
with the in-memory config store and local batch backend; wiring a Postgres DSN
and an installed/deployed `pg_llm_batch` client activates the KV/secret stores,
Expand Down
Loading
Loading