Skip to content
Open
Show file tree
Hide file tree
Changes from 22 commits
Commits
Show all changes
227 commits
Select commit Hold shift + click to select a range
9b0ddee
docs(doctoring): trace gateway failure and publisher permissions
seonghobae Sep 9, 2026
15b8f52
docs(research): verify stored paper versions and license declarations
seonghobae Sep 9, 2026
7ad34ff
docs(research): record package archive inspection evidence
seonghobae Sep 9, 2026
59a4c98
docs(analytics): distinguish active quality goal from legacy valuatio…
seonghobae Sep 9, 2026
29f417f
docs(research): address provenance review findings
seonghobae Sep 9, 2026
279f7e0
docs(research): trace joint speed and accuracy hypothesis
seonghobae Sep 9, 2026
74efc7b
docs(research): correct latent interaction attribution
seonghobae Sep 9, 2026
dccd37e
docs(research): define autonomous KPI acceptance and handoff
seonghobae Sep 9, 2026
2aa6ce7
docs(loop): record KPI baseline, #1075 closure evidence, #1079 blocker
Sep 9, 2026
edd54e4
docs(research): trace multilevel follow-up from local bibliography
seonghobae Sep 9, 2026
7acb444
docs(research): record scoring and partial visual evidence
seonghobae Sep 9, 2026
323f0eb
docs: retain release-cycle section ownership
seonghobae Sep 9, 2026
ab8a7ca
docs(research): record completed candidate scoring tests
seonghobae Sep 9, 2026
d7bba88
fix(persistence): roll back failed state replacements
seonghobae Sep 9, 2026
1beb72c
docs(persistence): record state-write rollback reproduction
seonghobae Sep 9, 2026
aa67418
docs: separate atomicity and release-cycle guidance
seonghobae Sep 9, 2026
f1abe1e
test(persistence): verify rollback survives reopening
seonghobae Sep 9, 2026
877d511
docs(gap): track persistence fix and skipped review evidence
seonghobae Sep 9, 2026
b92d108
docs(research): separate model difficulty from batch transport
seonghobae Sep 9, 2026
99f8d87
docs(research): distinguish mean inference timing from decision p95
seonghobae Sep 9, 2026
5333b25
docs(evidence): retain failed judge calls in reported denominator
seonghobae Sep 9, 2026
b031d3a
docs(research): fingerprint stored paper versions
seonghobae Sep 9, 2026
716e012
test(persistence): cover deferred commit rollback
seonghobae Sep 9, 2026
4316be8
docs(persistence): record full-suite and commit-failure evidence
seonghobae Sep 9, 2026
34bf2f3
docs(papers): remove contradictory redistribution assurance
seonghobae Sep 9, 2026
ef374de
docs(research): record license correction and pending visual check
seonghobae Sep 9, 2026
0ea2a58
docs(psychometrics): bound benchmark prior calibration claims
seonghobae Sep 9, 2026
479bfe7
docs(loop): record autonomous KPI scope, PR1108 verification, hourly-…
seonghobae Sep 9, 2026
fcf0047
docs(research): trace expanded-population validity proposal
seonghobae Sep 9, 2026
8b749ba
docs(research): reconcile six cross-document paper references
seonghobae Sep 9, 2026
9a9f1ab
test(research): detect unindexed arxiv references
seonghobae Sep 9, 2026
8839bfc
docs(research): record inventory guard and negative check
seonghobae Sep 9, 2026
632831d
docs(loop): verify PR1109 atomicity, recount 88 PRs, harden stacking …
seonghobae Sep 9, 2026
f4bc3b7
docs(research): record completed regressions and multimodal evidence …
seonghobae Sep 9, 2026
4e078c5
docs(ci): trace incomplete stack-owner model reviews
seonghobae Sep 9, 2026
42e7ea1
docs(research): bound LSIRM model-selection discrepancy
seonghobae Sep 9, 2026
1b31166
docs(loop): verify PR1109 integer-index head, recount 88 PRs, name qu…
seonghobae Sep 9, 2026
9e08f14
docs(research): verify LSIRM discrepancy in archived journal PDF
seonghobae Sep 9, 2026
5619d70
docs(loop): re-observe PR1109 fuzzing green, recount 88 PRs, add fail…
seonghobae Sep 9, 2026
9427c29
docs(research): trace multidimensional LSIRM reproduction boundary
seonghobae Sep 9, 2026
12957d6
docs(research): index DOI-only multidimensional LSIRM lead
seonghobae Sep 9, 2026
e271918
test(persistence): exercise durable stream retention explicitly
seonghobae Sep 9, 2026
b2c0993
test(persistence): verify commit visibility before durable return
seonghobae Sep 9, 2026
143a2eb
docs(kpi): record durable acknowledgement verification boundary
seonghobae Sep 9, 2026
fbb933c
merge: retain rollback repair with current KPI evidence base
seonghobae Sep 9, 2026
ddf087d
docs(kpi): specify validated admission and capacity rejection accounting
seonghobae Sep 9, 2026
035b58c
docs(evidence): record integrated rollback regression and KPI visual …
seonghobae Sep 9, 2026
85a0b2d
test: require durable initial decision receipt over HTTP
seonghobae Sep 9, 2026
1a510fa
fix(ci): run repository quality checks for stacked pull requests
seonghobae Sep 9, 2026
c11df64
merge: retain rollback delta and activate stacked quality checks
seonghobae Sep 9, 2026
6a44284
docs(ci): retain stacked quality trigger repair know-how
seonghobae Sep 9, 2026
d721e04
merge: preserve canonical stacked quality repair and consolidate cont…
seonghobae Sep 9, 2026
8de7144
docs(incident): preserve request attribution and replay safety bounda…
seonghobae Sep 9, 2026
2996cd3
docs(loop): adopt stacked-quality merge, verify PR1108 restack, diagn…
seonghobae Sep 9, 2026
129a665
Merge remote-tracking branch 'origin/autoresearch/20260909-kpi-loop' …
seonghobae Sep 9, 2026
dcefcdb
docs(ci): record terminal contract failure and integrated repair
seonghobae Sep 9, 2026
814b9ea
feat: connect opt-in Rust decision receipt to HTTP routing candidate
seonghobae Sep 9, 2026
ce6e029
test(research): include DOI-only sources in citation discovery
seonghobae Sep 9, 2026
920df5a
test: require persisted admission before chat and SSE dispatch
seonghobae Sep 9, 2026
962b875
docs: record bounded DOI register visual inspection
seonghobae Sep 9, 2026
df6eb4a
feat: retain admission denominator and expose incomplete receipt joins
seonghobae Sep 9, 2026
7734e89
docs: constrain joint outcome and response time research
seonghobae Sep 9, 2026
9b32582
docs: record terminal rollback CI repair evidence
seonghobae Sep 9, 2026
01ce903
test: verify race receipt identity and incomplete admission evidence
seonghobae Sep 9, 2026
12b3957
docs: track request denominator and dispatch boundary findings
seonghobae Sep 9, 2026
8769c82
docs: distinguish terminal pending verdicts from live checks
seonghobae Sep 9, 2026
b6aa2f2
docs: define request measurement identity and failure boundaries
seonghobae Sep 9, 2026
aeb66e9
fix: preserve one admission across embedding failover and bound recei…
seonghobae Sep 9, 2026
07957ee
Merge branch 'codex/state-save-rollback-20260909' into codex/decision…
seonghobae Sep 9, 2026
b3ba5bc
docs: record integrated receipt checkpoint and query plan evidence
seonghobae Sep 9, 2026
6e4f87a
fix: index receipt cohorts and distinguish repeated race decisions
seonghobae Sep 9, 2026
fe2db40
Merge branch 'codex/error-request-correlation-20260908' into codex/de…
seonghobae Sep 9, 2026
0ef13ed
docs: distinguish auxiliary dispatch from task routing latency
seonghobae Sep 9, 2026
2bd3fcd
docs: record indexed receipt and identity integration evidence
seonghobae Sep 9, 2026
05b512e
Merge commit '0ef13edc' into codex/decision-latency-receipts-20260909
seonghobae Sep 9, 2026
9c6d8cc
docs(loop): absorb PR1108 loop-merge, note PR1105 unhealed, add wrong…
seonghobae Sep 9, 2026
fd702ef
docs: record bounded task interval visual inspection
seonghobae Sep 9, 2026
8c4edbc
feat: separate auxiliary provider timing from task decisions
seonghobae Sep 9, 2026
ad67581
docs: track cache terminal and auxiliary export gaps
seonghobae Sep 9, 2026
0e6cd58
fix: export bounded auxiliary timing diagnostics
seonghobae Sep 9, 2026
30b7b3a
docs: record verified local literature discovery access
seonghobae Sep 9, 2026
21d1e88
fix: retain cache-hit admissions without provider timings
seonghobae Sep 9, 2026
b5f9a5a
docs: correct inherited publishing secret evidence
seonghobae Sep 9, 2026
3478b90
test: cover diagnostic cohort export cap boundaries
seonghobae Sep 9, 2026
cf5fdb4
docs(loop): note static PRs and fetch transient, add merge-delete rat…
seonghobae Sep 9, 2026
6ddddb8
fix: keep generated planning inside task decision interval
seonghobae Sep 9, 2026
9707a5e
fix: distinguish post-decision evidence embedding observations
seonghobae Sep 9, 2026
79658cb
docs: define per-invocation evidence timing phases
seonghobae Sep 9, 2026
3f42417
docs: bind package acceptance to canonical release owner
seonghobae Sep 9, 2026
bbe7eae
build: verify locked native receipts before CI acceptance
seonghobae Sep 9, 2026
fa620fb
docs: correct streaming admission hypothesis with HTTP evidence
seonghobae Sep 9, 2026
0b10b55
fix: preserve unmeasured streaming handler compatibility
seonghobae Sep 9, 2026
9467264
docs: preserve full regression failures and repair evidence
seonghobae Sep 9, 2026
bafbaf9
fix: admit streaming triage and retain typed selection failures
seonghobae Sep 9, 2026
2f3a086
Merge commit 'daa5e3636342644c49a037e705e2ecf85e642c71' into codex/de…
seonghobae Sep 9, 2026
3b6dd47
fix: finalize admitted request errors without erasing acknowledgements
seonghobae Sep 9, 2026
d3b9f12
docs: ground response process covariates in psychometric evidence
seonghobae Sep 9, 2026
a10ec0f
docs: distinguish completeness requirement from observed evidence
seonghobae Sep 9, 2026
53a9266
docs: record full receipt regression and saturated triage gap
seonghobae Sep 9, 2026
c734567
fix: reserve chat streaming capacity before classification
seonghobae Sep 9, 2026
3cd903a
docs: identify durable outcome linkage prerequisite for accuracy
seonghobae Sep 9, 2026
6b40db4
fix: retain durable workflow request origin identity
seonghobae Sep 9, 2026
4cf7fea
docs: record exact workflow linkage reproduction environment
seonghobae Sep 9, 2026
6098761
docs: record capacity regression and linkage test prerequisite
seonghobae Sep 9, 2026
af8d732
docs: repair locked full-suite environment prerequisite
seonghobae Sep 9, 2026
c226d30
docs: distinguish stacked review skip from approval
seonghobae Sep 9, 2026
c43a7d9
test: reproduce missed versioned arxiv citations
seonghobae Sep 9, 2026
304fd6f
fix: discover html and versioned arxiv citations
seonghobae Sep 9, 2026
9ac9315
fix: normalize versioned inventory identifiers
seonghobae Sep 9, 2026
8352d15
docs: preserve citation discovery red green evidence
seonghobae Sep 9, 2026
7636618
docs: qualify multilevel network IRT applicability
seonghobae Sep 9, 2026
6491591
docs: record direct multilevel IRT equation inspection
seonghobae Sep 9, 2026
2c578fa
docs: record completed linkage regression and research boundary
seonghobae Sep 9, 2026
387aa21
test: reproduce missing durable batch request lineage
seonghobae Sep 9, 2026
1546b9e
docs: link verified outcome package to stacked PR
seonghobae Sep 9, 2026
9b7f491
docs: record deferred batch lineage red evidence
seonghobae Sep 9, 2026
919c945
test: require explicit batch lineage write failure outcome
seonghobae Sep 9, 2026
9356ac6
fix: retain durable batch submission identity associations
seonghobae Sep 9, 2026
b565173
test: preserve repeated batch origins and standalone compatibility
seonghobae Sep 9, 2026
b5baf42
test: reject redundant remote batch registry write
seonghobae Sep 9, 2026
6223962
fix: avoid second remote registry write for lineage status
seonghobae Sep 9, 2026
f61f206
docs: qualify batch lineage event and failure evidence
seonghobae Sep 9, 2026
3773f0d
test: distinguish registry snapshots from durable lineage proof
seonghobae Sep 9, 2026
2c175b4
docs: record hosted receipt platform acceptance
seonghobae Sep 9, 2026
6966974
test: retain remote batch handle after registry failure
seonghobae Sep 9, 2026
d95ad53
fix: preserve applied batch outcome across registry failures
seonghobae Sep 9, 2026
369ea1e
fix: keep registry write result outside persisted snapshot
seonghobae Sep 9, 2026
18af294
test: distinguish absent and partially written batch registry
seonghobae Sep 9, 2026
b12981d
test: inspect canonical registry key in partial write probe
seonghobae Sep 9, 2026
cbbdaa2
docs: track batch lineage repair and recovery gap
seonghobae Sep 9, 2026
d71bcc0
test: reproduce missing owner-bound batch restart recovery
seonghobae Sep 9, 2026
6b40cce
fix: recover owner-bound remote batch descriptors from durable events
seonghobae Sep 9, 2026
e8d8c8a
fix: use existing attribution dimensions normalizer
seonghobae Sep 9, 2026
65b4160
test: reject unsafe batch recovery descriptors and result identities
seonghobae Sep 9, 2026
853e860
fix: validate batch recovery shape and preserve configured expiry
seonghobae Sep 9, 2026
b4abdd6
docs: record owner-bound batch restart recovery evidence
seonghobae Sep 9, 2026
aee00ac
docs: record batch recovery checkpoint and remaining outage checks
seonghobae Sep 9, 2026
e372bc5
test: reject inconsistent recovery identities and survive registry ou…
seonghobae Sep 9, 2026
7fb1a71
fix: isolate authorized batch recovery from unavailable registry
seonghobae Sep 9, 2026
83394af
test: bind recovery to deployment and reject malformed job fields
seonghobae Sep 9, 2026
f4d036b
test: recover backend metadata despite healthy coordinator handle
seonghobae Sep 9, 2026
4a09060
docs: retain recovery outage regressions and visual receipt
seonghobae Sep 9, 2026
6cb530a
fix: recover partial registry state and expose descriptor availability
seonghobae Sep 9, 2026
f674305
docs: record exact hosted outcome linkage validation
seonghobae Sep 9, 2026
42ad396
test: reject duplicate persisted items and verify recovery polling
seonghobae Sep 9, 2026
aaa9b13
fix: require exact persisted batch item cardinality
seonghobae Sep 9, 2026
17cfa61
test: retain applied remote submission after backend registry failure
seonghobae Sep 9, 2026
be090bd
docs: track pre-return submission and recovery lifetime gaps
seonghobae Sep 9, 2026
8ddfeb9
fix: return applied batch handle despite backend registry failure
seonghobae Sep 9, 2026
df638d6
test: preserve healthy active batch beyond recovery expiry
seonghobae Sep 9, 2026
a6b9485
fix: preserve healthy batch registry lifecycle during recovery
seonghobae Sep 9, 2026
1c8a1c1
docs: qualify Baker citation and parameter recovery metric
seonghobae Sep 9, 2026
3f98f92
docs: record batch recovery review repairs and deployment binding
seonghobae Sep 9, 2026
08a660d
test: prevent healthy registry from bypassing deployment binding
seonghobae Sep 9, 2026
cd38d9c
fix: validate deployment binding before healthy batch fast path
seonghobae Sep 9, 2026
2d68f70
test: preserve explicitly unbound legacy batch compatibility
seonghobae Sep 9, 2026
1c405a1
docs: retain submission lifecycle and target binding regressions
seonghobae Sep 9, 2026
c06615a
test: preserve endpoint binding on healthy batch metadata
seonghobae Sep 9, 2026
4506675
fix: enforce exact endpoint on bound batch metadata
seonghobae Sep 9, 2026
73404f8
docs: freeze reviewed batch recovery acceptance scope
seonghobae Sep 9, 2026
84861c1
docs: record reviewed batch freeze and source copyright inspection
seonghobae Sep 9, 2026
cd5f4f6
docs: separate synthetic recovery fixture from routing evidence
seonghobae Sep 9, 2026
c17d2e0
docs: record installed recovery proof and bounded IRT conditions
seonghobae Sep 9, 2026
4c81ead
docs: capture installed request-outcome export gap
seonghobae Sep 9, 2026
2c07529
docs: record export regression and historical evidence constraint
seonghobae Sep 9, 2026
9b9048c
docs: record batch recovery full verification and stacked PR
seonghobae Sep 9, 2026
a50c3bc
docs: pin Fugu source and separate learned selection from role fixtures
seonghobae Sep 9, 2026
5a84d57
docs: distinguish Fugu training reward from correctness
seonghobae Sep 9, 2026
06ff237
docs: reconcile export review and hosted batch evidence
seonghobae Sep 9, 2026
204e306
docs: record exact-head hosted batch acceptance
seonghobae Sep 9, 2026
f8e77ba
docs: map Fugu evidence to current optimizer boundaries
seonghobae Sep 9, 2026
880a800
docs: track Fugu memory and comparison constraints
seonghobae Sep 9, 2026
1c25fbb
docs: track invalid optimizer score repair evidence
seonghobae Sep 9, 2026
89a79d0
docs: pin Fugu evaluation protocol caveats
seonghobae Sep 9, 2026
ee4d7b0
docs: separate adaptive experiments from independent evidence
seonghobae Sep 9, 2026
868ca4f
docs: record installed optimizer score validation
seonghobae Sep 9, 2026
82892ed
Merge pull request #1112 from ContextualWisdomLab/codex/decision-late…
seonghobae Sep 9, 2026
5a9d8fb
Merge pull request #1113 from ContextualWisdomLab/codex/outcome-reque…
seonghobae Sep 9, 2026
53dcd1b
Merge pull request #1115 from ContextualWisdomLab/codex/batch-request…
seonghobae Sep 9, 2026
e306525
Merge branch 'codex/state-save-rollback-20260909' (PR #1108 head 53dc…
seonghobae Sep 10, 2026
52fd0da
merge: reconcile current main with retained KPI research stack
seonghobae Sep 12, 2026
d8a0a72
docs(kpi): record preserved stack and missing acknowledgement failure
seonghobae Sep 12, 2026
36af4a5
test(receipts): await durable request finalization before export
seonghobae Sep 12, 2026
81ad64c
docs(receipts): explain finalization race and verified synchronization
seonghobae Sep 12, 2026
0fd408a
docs: record terminal research stack verification
seonghobae Sep 12, 2026
be60791
docs: align product gap with integrated test and visual evidence
seonghobae Sep 12, 2026
dcaf2b2
docs: qualify LaRT measurement evidence and owner experiment
seonghobae Sep 12, 2026
9a49cb8
fix(research): register LaRT in canonical paper inventory
seonghobae Sep 12, 2026
f814e06
merge: retain research delta on repaired KPI parent
seonghobae Sep 12, 2026
de01e9e
test(research): distinguish marginal and conditional trait uncertainty
seonghobae Sep 12, 2026
ff590dc
docs: record LaRT unit and scoped visual verification
seonghobae Sep 12, 2026
60ee94c
docs: distinguish intentional draft skips from central dispatch evidence
seonghobae Sep 12, 2026
4359401
docs: track nonlinear response dependence research gap
seonghobae Sep 12, 2026
61380f2
docs: separate dry-run readiness from observed KPI evidence
seonghobae Sep 12, 2026
c8b358a
test: reject production evidence labels on dry-run reports
seonghobae Sep 12, 2026
383b4a1
fix: label dry-run reports as synthetic diagnostics
seonghobae Sep 12, 2026
07b95c9
docs: record synthetic benchmark classification repair
seonghobae Sep 12, 2026
eeefaca
docs: record full and installed synthetic classification verification
seonghobae Sep 12, 2026
c4e8343
test(receipts): reproduce request cache aggregation failures
seonghobae Sep 12, 2026
0ddb5f1
fix(receipts): finalize cache reuse at request boundary
seonghobae Sep 12, 2026
4bc9604
test(receipts): verify cache context cleanup between requests
seonghobae Sep 12, 2026
36ade58
docs: record cache aggregation acceptance and review handoff gap
seonghobae Sep 12, 2026
9863638
docs: retain bounded cache repair visual inspection
seonghobae Sep 12, 2026
7e2fc52
docs: record free review transport source comparison
seonghobae Sep 12, 2026
1b8cb34
docs: qualify registry and release evidence scope
seonghobae Sep 12, 2026
d91c102
docs(review): record exact protected merge startup proof
seonghobae Sep 12, 2026
0612344
docs(review): record rendered transport evidence inspection
seonghobae Sep 12, 2026
3cf8487
docs(gap): track existing review repair adoption boundary
seonghobae Sep 12, 2026
a73737e
docs(research): qualify LaRT split and separation assumptions
seonghobae Sep 12, 2026
369f891
docs(research): quantify pinned model-prompt split overlap
seonghobae Sep 12, 2026
85fabb1
docs(gap): gate retrospective calibration on verified split identities
seonghobae Sep 12, 2026
a71ae99
docs(research): record visual inspection of split audit
seonghobae Sep 12, 2026
1ffd5d4
docs(research): link installed owner partition rejection evidence
seonghobae Sep 12, 2026
843f941
docs(research): record installed gate visual verification
seonghobae Sep 12, 2026
f3a81f8
docs(research): trace candidate combined item identities to upstream …
seonghobae Sep 12, 2026
c2d98e8
docs(research): record exact combined token-count offset audit
seonghobae Sep 12, 2026
75c9d29
docs(research): trace predictive preprocessing offset and remaining a…
seonghobae Sep 12, 2026
b4887c2
docs(research): record predictive preprocessing visual receipt
seonghobae Sep 12, 2026
14a6a94
docs: transfer preprocessing follow-ups to verified successor with re…
seonghobae Sep 12, 2026
de21ffd
docs: reconcile multilevel IRT citation with verified DOI
seonghobae Sep 12, 2026
aeca0df
docs: prioritize observed quality and decision latency KPIs
seonghobae Sep 12, 2026
b0844bd
docs: register bounded conditional independence research
seonghobae Sep 12, 2026
8f7d583
docs(research): distinguish fitted residual diagnostics from KPI evid…
seonghobae Sep 12, 2026
e9d51ac
docs(research): preserve separate uncertainty check heading
seonghobae Sep 12, 2026
85827d2
docs(research): link residual diagnostic owner acceptance gap
seonghobae Sep 12, 2026
9788cb4
docs(gap): distinguish KPI field declarations from runtime evidence
seonghobae Sep 12, 2026
5a12767
docs(gap): connect request latency candidate to acceptance evidence
seonghobae Sep 12, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
11 changes: 11 additions & 0 deletions .github/opencode/hourly-loop-prompt.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,6 +9,17 @@ completion, explicit user/operator cancellation, and the hosting platform's own
execution contract are the only time-based termination authorities available to
this agent. Never post intermediate progress reports.

PR 0 only via merge or verified-successor full-delta inheritance. Never
force-push, never close without evidence (user-explicit, no valid delta,
malicious change, or verified complete inheritance only). Single-writer
deltas are integrated, never discarded. Before and after long runs record
`git rev-parse HEAD` and `git diff --stat origin/main...HEAD --
contextual_orchestrator/ tests/`; docs-only drift does not invalidate code
evidence. Prefer isolated worktrees, preserve live execution handles, and
never rebase or push another session's branch. Synthetic recovery is unit
evidence only; customer accuracy and decision latency require observed
outcomes with declared denominators and uncertainty.

Before changing code, read `docs/product_planning.md`, the applicable PRD in
`docs/model-group-product-technical-spec.md`, and
`docs/product-technical-gap-baseline.md`. Treat current files and exact GitHub
Expand Down
9 changes: 9 additions & 0 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -3,6 +3,15 @@
Cross-agent conventions for `contextual-orchestrator`, readable by any coding
agent (Claude, Codex, Cursor, opencode, …). Keep this file tool-agnostic.

## Autonomous research handoff

Read [the KPI runbook](docs/doctoring/autonomous_kpi_runbook.md) before numerical
experiments and update its verified evidence before handoff. Choose KPI scope
autonomously under `docs/analytics_spec.md`. Preserve live execution handles;
high host load and silent numerical work are not proof of deadlock. Synthetic
recovery is unit evidence, not customer accuracy. Break owner/consumer release
cycles with isolated exact-revision contracts, never production source copies.

<!-- BEGIN cwl-agent-guidance -->
## Agent guidance (CWL governance)

Expand Down
8 changes: 8 additions & 0 deletions CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,6 +4,14 @@ This file provides guidance to Claude Code (claude.ai/code) when working with co

## Read AGENTS.md first

For autonomous experiments, also read and maintain
[the single KPI runbook](docs/doctoring/autonomous_kpi_runbook.md).
It records commands, environment limits, failed interpretations, and evidence
boundaries. Do not restart a live numerical test because polling is silent or
claim customer improvement from synthetic recovery. Use the independently
verifiable owner/consumer sequence recorded there instead of waiting for all
foundation releases before developing a port.

Equivalent model-group endpoints may race only through the normalized, explicit
endpoint-equivalence contract. Preserve modality validation, bounded concurrency,
deadline, cancellation/drain provenance, and honest duplicate-cost evidence.
Expand Down
37 changes: 17 additions & 20 deletions contextual_orchestrator/benchmark_priors.py
Original file line number Diff line number Diff line change
@@ -1,40 +1,37 @@
"""Benchmark-quality priors for the model-group Beta ledgers.

Two public measurements feed each member's prior success probability:
LMSYS Chatbot Arena Elo (Bradley–Terry rating scale) and Artificial
Analysis' Quality Index. Both are *published measurements*; this module
never invents a numeric weight. Everything else is derived from either

1. those measurements themselves,
2. an existing repository constant (``model_group``'s Laplace prior
budget), or
3. arithmetic over the items above.
Legacy heuristic priors combine shipped Chatbot Arena ratings and Artificial
Analysis Quality Index values. Their snapshot provenance has not been verified
against archived source data. The transform below is deterministic but is not a
calibrated probability of task success or an identified psychometric estimate.

Derivation contract (auditable, deterministic):

- Each shipped rating ``r_i`` is centered on the median of the shipped
set and scaled by the set's own median absolute deviation (MAD),
giving ``z_i = (r_i - median) / MAD_i``.
- The two instruments are averaged after normalization (they measure
overlapping-but-distinct constructs; equal weight is the maximum
entropy choice across exactly two sources, not a tuned parameter).
- ``p_hat = logistic(z)`` is then a posterior-style membership value in
``(0, 1)`` measured from ratings alone.
- The two instruments are averaged after normalization. Equal weighting is
an implementation choice, not a fitted weight justified by these references.
- ``p_hat = logistic(z)`` maps the composite to ``(0, 1)``. That range alone
does not establish posterior or predictive calibration.
- The prior is *mass preserving*: ``(alpha0, beta0)`` splits the exact
unobserved-evidence budget that ``model_group`` already spends on any
unknown member (its Laplace counts), so a known member never receives
more evidence than an unknown one — it only receives that identical
budget distributed according to measurement instead of uniformly.

Failure denominator: members absent from every shipped instrument fall
back to the unchanged repository Laplace prior.
Members missing from either instrument fall back to the existing Laplace prior.
Neither the fixed evidence budget nor that fallback proves validity. Replacing
this legacy behavior requires observed-task calibration and release evidence;
do not treat these values as customer accuracy evidence.

References (APA 7th):
Bradley, R. A., & Terry, M. E. (1952). Rank analysis of incomplete
block designs: I. The method of paired comparisons. *Biometrika,
39*(3/4), 324–345. https://doi.org/10.1093/biomet/39.3-4.324
Chiang, W., Zheng, L., Ma, Z., Li, Y., Sheng, Z., Wu, X., ... Zhang,
H. (2024). *Chatbot Arena: An open platform for evaluating LLMs
Chiang, W.-L., Zheng, L., Sheng, Y., Angelopoulos, A. N., Li, T.,
Li, D., Zhang, H., Zhu, B., Jordan, M. I., Gonzalez, J. E., &
Stoica, I. (2024). *Chatbot Arena: An open platform for evaluating LLMs
by human preference* [Preprint]. arXiv.
https://doi.org/10.48550/arXiv.2403.04132
"""
Expand Down Expand Up @@ -123,7 +120,7 @@ def _normalized_membership(name: str) -> float | None:


def measured_quality_probability(member_id: str) -> float | None:
"""Return the measurement-derived prior success probability, if known."""
"""Return the legacy heuristic prior fraction, if both scores are known."""
lowered = member_id.lower()
for key in _ARENA_ELO:
if key in lowered:
Expand All @@ -134,7 +131,7 @@ def measured_quality_probability(member_id: str) -> float | None:


def resolve_quality_prior(member_id: str) -> tuple[float, float]:
"""Resolve the benchmark-measured ``(alpha, beta)`` prior for one member.
"""Resolve the legacy benchmark-derived ``(alpha, beta)`` prior.

Unknown members receive the repository's unchanged Laplace pair, so
behaviour for unmeasured identifiers is bit-for-bit the pre-existing
Expand Down
32 changes: 31 additions & 1 deletion docs/analytics_spec.md
Original file line number Diff line number Diff line change
Expand Up @@ -140,9 +140,39 @@ never present it as an all-request guarantee or encode missing durations as zero
Declare the observation window, quantile method, uncertainty method, and workload
before comparing policies. A faster failed decision is not a quality improvement.

### Autonomous experiment targets

These are engineering acceptance targets selected on 2026-09-09, not measured
results or literature-derived constants. Do not ask the user to choose their
scope. Baseline and candidate must use the same declared population, endpoint
mix, task rubric, model versions, resource budget, and failure accounting.

| Outcome | Target | Guardrail |
| --- | --- | --- |
| Delivered-correct fraction over all accepted requests | At least +1 percentage point versus baseline, with a 95% confidence interval for the difference wholly above zero. | Independently adjudicated observed outcomes; paired or randomized design declared before evaluation; no silent exclusion of failed delivery. |
| Routing-decision p95 | At most 20 ms and at least 10% lower than baseline, with the 95% interval for the candidate/baseline ratio wholly below 1. | Include selection and durable acknowledgement; preserve workload and failure accounting. This is not the full-page latency SLO. |
| Numerical parameter recovery | No regression in family-wise aligned RMSE under the declared numerical tolerance. | Known-truth unit tests only; never substitute for observed customer outcomes. |

The end target requires both customer accuracy and decision-latency criteria.
Intermediate changes may advance one while preserving the other's established
baseline; they must not be labelled completion of both. If non-regression cannot
be established, keep the candidate experimental and leave production unchanged.
Use a fresh holdout for confirmation after adaptive experiment selection, or a
predeclared sequential inference procedure; repeated inspection of an ordinary
95% interval is not a stopping rule. Determine sample size from pilot variance,
the +1-point effect target, and declared power before confirmation, retaining
task/model/time clustering. Do not reduce sample size after seeing results.

Execution and evidence handoff: [autonomous KPI runbook](doctoring/autonomous_kpi_runbook.md).

## Commercial Due-Diligence KPIs

These metrics support the KRW 2,000,000,000 commercial-readiness review. The
The active goal uses a USD 20,000,000,000 sale-quality ambition. This is an
aspirational quality target, not a measured valuation or a signed transaction.
The KRW 2,000,000,000 references in the legacy metrics below describe the
earlier commercial-readiness review; they must not substitute for the active
goal or be treated as its currency conversion. Accuracy and decision-latency
acceptance still require the observed evidence specified above. The
`evidence_type` column is mandatory so measured local evidence is never mixed
with proposed production targets.

Expand Down
9 changes: 8 additions & 1 deletion docs/benchmarks/2026-08-11-polytomous-llm-judge.md
Original file line number Diff line number Diff line change
Expand Up @@ -3,6 +3,12 @@
Status: exploratory integration evidence; not a claim that the local judge is
unbiased or that this single case is sufficient for IRT estimation.

Evidence classification (2026-09-09 audit): these historical model calls use
constructed release-plan cases. They are not a probability sample of customer
requests and must not supply the observed delivered-correct KPI or parameter
recovery evidence. Tables below are historical reports, not freshly reproduced
raw-artifact verification.

## Setup

The path under test was:
Expand Down Expand Up @@ -137,7 +143,8 @@ parser rejected the model response; it was never repaired or accepted.
| liked | `invalid` / `0.7500; yes` / `0.8333; yes` | `0.5000; no` / `0.0000; no` / `invalid` |
| disliked | `0.5000; no` / `0.7500; yes` / `0.8333; yes` | `invalid` / `invalid` / `invalid` |

The good plan parsed in 11/18 calls and was accepted in 5/11 parsed calls;
The good plan parsed in 11/18 calls and was accepted in 5/11 parsed calls
(45.45%, conditional on parsing), or 5/18 of all attempted calls (27.78%);
the seven failures were five invalid JSON responses, one out-of-range category,
and one non-monotone threshold vector. The unsafe plan parsed in all 18 calls,
scored `0.0000` in every case, and was accepted zero times. Direct judging
Expand Down
Loading