Skip to content
Merged
Show file tree
Hide file tree
Changes from 2 commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 8 additions & 2 deletions .github/workflows/verify-agent-bundles.yml
Original file line number Diff line number Diff line change
Expand Up @@ -37,7 +37,7 @@ jobs:
- name: Install the parser
# No PyPI release matches the fork build the server runs, so this gate
# approximates it. ecr-server.yml parses with what production has.
run: pip install --disable-pip-version-check 'omnigent==0.9.0'
run: pip install --disable-pip-version-check 'omnigent==0.16.0'

- name: Parse every bundle and assert its contract
run: |
Expand All @@ -55,7 +55,7 @@ jobs:
# A scout carries no method and takes none from the host: the driver's
# per-run prompt is its whole instruction, so a discovered skill would
# be capability it was never asked to use.
'xreview-scout-codex': {'filter': 'none', 'skills': 0},
'xreview-scout-codex': {'filter': 'none', 'skills': 0, 'effort': 'high'},
# Same contract on a different harness. Listed here because a bundle
# with no entry fails below, and the deployment cannot reach its
# harness yet: this gate is the only thing asserting it stays parsable
Expand Down Expand Up @@ -102,6 +102,12 @@ jobs:
print(f'FAIL {d.name}: non-{pref} skills vendored in: {stray}')
failed = True

# A parser too old to read the field reports None, which fails here too.
effort = getattr(spec.executor, 'reasoning_effort', None)
if 'effort' in exp and effort != exp['effort']:
print(f'FAIL {d.name}: reasoning effort is {effort!r}, expected {exp["effort"]!r}')
failed = True
Comment thread
cursor[bot] marked this conversation as resolved.

# With `none`, host discovery must return nothing even from a directory
# that has host skills to offer.
if spec.skills_filter == 'none':
Expand Down
14 changes: 5 additions & 9 deletions agents/xreview-scout-codex/config.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -28,15 +28,11 @@ executor:
# vendor-neutral gateway, read live off a process that was routed at
# api.openai.com. Misreading it produced three wrong diagnoses.
#
# gpt-5.6-sol is what ai-review's codex pass resolves to: that workflow pins no
# model and takes openai/codex-action's default, measured on a passing run.
# Matching it is the point — a scout on a different model than the reference
# reviewer is not the second opinion we mean to reproduce.
#
# Dotted, not dashed: codex spells it gpt-5.6-sol where the catalog writes
# gpt-5-6-sol. This value goes verbatim to `codex -c model=`, and the dotted
# form is the one the API accepts.
model: gpt-5.6-sol
# gpt-6.1-sol is the Codex CLI's default model. Dotted, not dashed: this value
# goes verbatim to `codex -c model=`, and the dotted form is the one the API
# accepts.
model: gpt-6.1-sol
reasoning_effort: high
Comment thread
seidroid[bot] marked this conversation as resolved.
config:
# codex, and that is the point. The bundle is what fixes the harness,
# so this is the one field that makes this a genuinely different reader
Expand Down
Loading