Skip to content
Merged
Show file tree
Hide file tree
Changes from 2 commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .github/workflows/opencode-review-dispatch.yml
Original file line number Diff line number Diff line change
Expand Up @@ -3990,7 +3990,7 @@ jobs:
"npm": "@ai-sdk/openai-compatible",
"name": "Contextual Orchestrator",
"options": {
"baseURL": "{env:CONTEXTUAL_ORCHESTRATOR_BASE_URL}",
"baseURL": "{env:CONTEXTUAL_ORCHESTRATOR_BASE_URL}/v1",
"apiKey": "{env:CONTEXTUAL_ORCHESTRATOR_TOKEN}"
},
"models": {
Expand Down
2 changes: 1 addition & 1 deletion .github/workflows/pr-review-autofix.yml
Original file line number Diff line number Diff line change
Expand Up @@ -364,7 +364,7 @@ jobs:
"npm": "@ai-sdk/openai-compatible",
"name": "Contextual Orchestrator",
"options": {
"baseURL": "{env:CONTEXTUAL_ORCHESTRATOR_BASE_URL}",
"baseURL": "{env:CONTEXTUAL_ORCHESTRATOR_BASE_URL}/v1",
"apiKey": "{env:CONTEXTUAL_ORCHESTRATOR_TOKEN}"
},
"models": {
Expand Down
13 changes: 13 additions & 0 deletions .github/workflows/strix.yml
Original file line number Diff line number Diff line change
Expand Up @@ -976,6 +976,12 @@ jobs:
# Defined before the gate loop so the bounded retry decision below
# can classify outcomes without duplicating the patterns later.
backend_unavailable_signal='STRIX_PROVIDER_UNAVAILABLE|RateLimitError|Too many requests\. For more on scraping GitHub|exceeded your current quota|insufficient_quota|billing details|"status"[[:space:]]*:[[:space:]]*"RESOURCE_EXHAUSTED"|tokens_limit_reached|Request body too large|Max size:[[:space:]]*[0-9]+[[:space:]]+tokens|Error code:[[:space:]]*500[^[:cntrl:]]*internal_error|Error code:[[:space:]]*413|LLM CONNECTION FAILED|Could not establish connection to the language model|LLM warm-up failed|Configured model and fallback models were unavailable|Configured Vertex model and fallback models were unavailable|emitted provider infrastructure or failure-signal output|before provider infrastructure failure|litellm(\.exceptions)?\.NotFoundError[^[:cntrl:]]*Nvidia_nimException[^[:cntrl:]]*Error code:[[:space:]]*404|Error during penetration test: loginAsGuest failed after [0-9]+ attempts: curl exit 7: curl: \(7\) Failed to connect to 127\.0\.0\.1 port 48080'
# Scanner tooling breakage (Caido GraphQL query/cursor errors) is
# not a provider outcome. It can occur while providers are healthy
# and it can coexist with genuine provider rate limits, so it gets
# its own typed notice instead of being folded into the provider
# verdict. See docs/doctoring/review-failure-taxonomy.md.
tooling_error_signal='Invalid HTTPQL query|Failed to parse cursor|TransportQueryError|caido_sdk_client\.errors'
model_behavior_error_signal='(^|[^A-Za-z0-9_])(agents|pydantic_ai|strix)(\.[A-Za-z_][A-Za-z0-9_]*)*\.ModelBehaviorError([^A-Za-z0-9_]|$)'
# Any evidence that a vulnerability was actually reported. Its presence
# forces a hard failure so real findings are NEVER downgraded. Keep the
Expand Down Expand Up @@ -1029,6 +1035,13 @@ jobs:
# Classify provider/backend exhaustion only when no vulnerability
# finding was emitted. Classification improves diagnosis; it never
# converts an incomplete scan into passing security evidence.
# Report scanner tooling breakage on its own, whatever the provider
# verdict turns out to be. This never changes the exit code: an
# incomplete scan stays non-passing either way.
if grep -Eq "$tooling_error_signal" "$strix_neutralization_scope_log"; then
echo "::error title=STRIX_TOOLING_ERROR::Strix scanner tooling failed (Caido GraphQL query or cursor error). This is a scanner defect, not a provider outage; a provider notice may also follow. See the strix-reports artifact and run log."
fi

if ( grep -Eiq "$backend_unavailable_signal" "$strix_neutralization_scope_log" \
|| grep -Eq "$model_behavior_error_signal" "$strix_neutralization_scope_log" ) \
&& ! grep -Eiq "$reported_vulnerability_signal" "$strix_neutralization_scope_log"; then
Expand Down
51 changes: 51 additions & 0 deletions docs/doctoring/review-failure-taxonomy.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,51 @@
# Review failure taxonomy: gateway routing, scanner tooling, dispatch admission

On 2026-09-21 the central review pipeline looked like a provider outage. It was not.
Three unrelated failures were being read as one.

## What the hosted evidence showed

Every run reached the vendored contextual-orchestrator sidecar and every run reported
`provider secrets present: 5 of 5`, including runs that predate any credential change.
The gateway answered its own preflight with `status: ready` and `finish_reason: stop`.
Provider credentials were never the blocker.

## 1. OpenCode: gateway routing, reported as `Error: not found`

The sidecar exports `CONTEXTUAL_ORCHESTRATOR_BASE_URL` as a bare `scheme://host:port`.
Noema and Strix append `/v1/chat/completions` themselves. OpenCode's
`@ai-sdk/openai-compatible` provider appends only `/chat/completions`, so it posted to an
unprefixed path. The gateway serves `/v1/chat/completions` and answers anything else with
`route_not_found`, whose message is the bare string `not found` — which OpenCode printed
verbatim as `Error: not found`, half a second after its banner had already resolved the
agent and model.

Reproduced locally against a stub gateway with the installed OpenCode CLI: an unprefixed
`baseURL` produced `POST /chat/completions`, and `{env:CONTEXTUAL_ORCHESTRATOR_BASE_URL}/v1`
produced `POST /v1/chat/completions`. The `/v1` belongs in the OpenCode provider options,
never in the sidecar export — moving it there would double-prefix Noema and Strix.

The banner is the tell: once `> <agent> · <model>` has printed, agent and model already
resolved, so a later `not found` is a transport answer, not configuration lookup.

## 2. Strix: scanner tooling, previously folded into the provider verdict

A failing scan emitted Caido GraphQL errors (`Invalid HTTPQL query`, `Failed to parse
cursor`, `TransportQueryError`) while the gateway was healthy. Genuine provider rate
limits appeared in the same log, so the single `STRIX_PROVIDER_UNAVAILABLE` notice was not
wrong — it was incomplete, and it hid a scanner defect behind an infrastructure label.
`strix.yml` now emits `STRIX_TOOLING_ERROR` on its own whenever a tooling signature
appears. Both notices can appear together. Neither changes the exit code: an incomplete
scan stays non-passing.

## 3. Dispatch admission: never a gateway outcome

`repository_dispatch authorization rejected` and `repository_dispatch metadata does not
match the live pull request` fire before the sidecar is provisioned. Counting them as
review-pipeline outages inflates the apparent provider failure rate.

## Rule

Attribute a review failure to the provider only after the gateway request itself failed.
Name the gateway's served path, the scanner's own errors, and admission gates as separate
classes. `tests/test_review_failure_taxonomy_contract.py` pins all three.
7 changes: 5 additions & 2 deletions opencode.jsonc
Original file line number Diff line number Diff line change
Expand Up @@ -294,12 +294,15 @@
// routes prioritized by scripts/ci/zdr_policy.py. Requires
// CONTEXTUAL_ORCHESTRATOR_BASE_URL and CONTEXTUAL_ORCHESTRATOR_TOKEN,
// which scripts/ci/contextual_orchestrator_review_sidecar.sh provisions on
// each runner before OpenCode starts.
// each runner before OpenCode starts. The sidecar exports a bare
// scheme://host:port, while the OpenAI-compatible provider appends only
// `/chat/completions`, so the `/v1` prefix belongs here. Without it the
// gateway answers route_not_found and OpenCode prints `Error: not found`.
"contextual-orchestrator": {
"npm": "@ai-sdk/openai-compatible",
"name": "Contextual Orchestrator",
"options": {
"baseURL": "{env:CONTEXTUAL_ORCHESTRATOR_BASE_URL}",
"baseURL": "{env:CONTEXTUAL_ORCHESTRATOR_BASE_URL}/v1",
"apiKey": "{env:CONTEXTUAL_ORCHESTRATOR_TOKEN}"
},
"models": {
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -437,7 +437,7 @@ def test_opencode_config_defaults_to_the_contextual_gateway() -> None:
assert f'"model": "{GATEWAY_MODEL}"' in config
assert f'"small_model": "{GATEWAY_MODEL}"' in config
assert '"enabled_providers": ["contextual-orchestrator"' in config
assert '"baseURL": "{env:CONTEXTUAL_ORCHESTRATOR_BASE_URL}"' in config
assert '"baseURL": "{env:CONTEXTUAL_ORCHESTRATOR_BASE_URL}/v1"' in config
assert '"apiKey": "{env:CONTEXTUAL_ORCHESTRATOR_TOKEN}"' in config
assert '"orchestrator/free": {' in config

Expand Down
4 changes: 2 additions & 2 deletions tests/test_pr_review_autofix_nvidia_nim_contract.py
Original file line number Diff line number Diff line change
Expand Up @@ -17,7 +17,7 @@
DOCTORING_RECORD = Path("docs/doctoring/hourly-nvidia-nim-autofix.md")
CHANGELOG = Path("CHANGELOG.md")
REVIEW_DISPATCH_WORKFLOW = Path(".github/workflows/opencode-review-dispatch.yml")
REVIEW_DISPATCH_BLOB_SHA = "cbc8d214394c4b7acbe82ce7fba11fd073b91c98"
REVIEW_DISPATCH_BLOB_SHA = "376e592b1f408db96ea14c935ca641da3e2162ac"


def _workflow_text(path: Path) -> str:
Expand All @@ -44,7 +44,7 @@ def test_scheduled_autofix_routes_through_contextual_orchestrator() -> None:
'"orchestrator/free": {',
'"reasoningEffort": "high"',
'"npm": "@ai-sdk/openai-compatible"',
'"baseURL": "{env:CONTEXTUAL_ORCHESTRATOR_BASE_URL}"',
'"baseURL": "{env:CONTEXTUAL_ORCHESTRATOR_BASE_URL}/v1"',
'"apiKey": "{env:CONTEXTUAL_ORCHESTRATOR_TOKEN}"',
"contextual_orchestrator_review_sidecar.sh",
"BYTEZ_API_KEY: ${{ secrets.BYTEZ_API_KEY }}",
Expand Down
130 changes: 130 additions & 0 deletions tests/test_review_failure_taxonomy_contract.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,130 @@
"""Contract tests separating gateway routing, tooling, and dispatch failures.

Three distinct failure classes were collapsed into one "provider unavailable"
reading during the 2026-09-21 review outage:

* OpenCode reached the vendored gateway but posted to an unprefixed path,
so the gateway answered ``route_not_found`` and OpenCode surfaced its
message verbatim as ``Error: not found``. The provider credentials were
healthy throughout.
* Strix emitted Caido GraphQL tooling errors alongside genuine provider
rate limits, and the workflow reported only the provider class.
* OpenCode Review Dispatch rejected a repository_dispatch before the
sidecar existed, which is neither a gateway nor a provider outcome.

These contracts keep each class independently observable.
"""

from __future__ import annotations

import json
from pathlib import Path
import re

_ORG_REPO_ROOT = Path(__file__).resolve().parents[1]

AUTOFIX_WORKFLOW = _ORG_REPO_ROOT / ".github/workflows/pr-review-autofix.yml"
OPENCODE_DISPATCH_WORKFLOW = (
_ORG_REPO_ROOT / ".github/workflows/opencode-review-dispatch.yml"
)
STRIX_WORKFLOW = _ORG_REPO_ROOT / ".github/workflows/strix.yml"
OPENCODE_CONFIG = _ORG_REPO_ROOT / "opencode.jsonc"

GATEWAY_SERVED_CHAT_PATH = "/v1/chat/completions"
EXPECTED_BASE_URL = "{env:CONTEXTUAL_ORCHESTRATOR_BASE_URL}/v1"


def _read(path: Path) -> str:
"""Return one tracked contract file as UTF-8 text."""
return path.read_text(encoding="utf-8")


def _generated_opencode_configs(workflow_text: str) -> list[dict]:
"""Return every ``jq -n`` generated OpenCode config in one workflow."""
matches = re.findall(
r"jq -n(?:[^']*)'(\{.*?\})' >\"\$\{[A-Z_]+\}/opencode\.jsonc\"",
workflow_text,
re.DOTALL,
)
return [json.loads(match) for match in matches]


def _strip_jsonc_comments(text: str) -> str:
"""Drop ``//`` line comments so a JSONC config parses as JSON."""
return "\n".join(
line for line in text.splitlines() if not line.lstrip().startswith("//")
)


def test_autofix_opencode_provider_targets_the_served_gateway_path() -> None:
"""The autofix provider baseURL must resolve to the gateway's /v1 routes."""
configs = _generated_opencode_configs(_read(AUTOFIX_WORKFLOW))
assert configs, "no generated OpenCode config found in the autofix workflow"
for config in configs:
options = config["provider"]["contextual-orchestrator"]["options"]
assert options["baseURL"] == EXPECTED_BASE_URL


def test_dispatch_opencode_provider_targets_the_served_gateway_path() -> None:
"""The dispatch gateway overlay must carry the same /v1 prefix.

The dispatch workflow writes a provider-free base config and then layers
the gateway provider on with a second ``jq`` filter, so the assertion is
on every ``baseURL`` the workflow binds for that provider.
"""
workflow = _read(OPENCODE_DISPATCH_WORKFLOW)
bound = re.findall(
r'"baseURL": "(\{env:CONTEXTUAL_ORCHESTRATOR_BASE_URL\}[^"]*)"', workflow
)
assert bound, "the dispatch workflow binds no gateway baseURL"
assert set(bound) == {EXPECTED_BASE_URL}


def test_repository_opencode_config_targets_the_served_gateway_path() -> None:
"""The tracked reviewer config must not drop the gateway's /v1 prefix."""
config = json.loads(_strip_jsonc_comments(_read(OPENCODE_CONFIG)))
options = config["provider"]["contextual-orchestrator"]["options"]
assert options["baseURL"] == EXPECTED_BASE_URL


def test_strix_reports_caido_tooling_failures_as_their_own_class() -> None:
"""Caido GraphQL breakage must not be reported only as a provider outage."""
workflow = _read(STRIX_WORKFLOW)
assert "tooling_error_signal=" in workflow
for signature in (
"Invalid HTTPQL query",
"Failed to parse cursor",
"TransportQueryError",
):
assert signature in workflow
assert "STRIX_TOOLING_ERROR" in workflow
tooling_index = workflow.index("STRIX_TOOLING_ERROR")
provider_index = workflow.index("STRIX_PROVIDER_UNAVAILABLE::Strix could not")
assert tooling_index < provider_index, (
"the tooling class must be emitted before the provider verdict so a "
"scanner defect is never filed only as a provider outage"
)


def test_dispatch_admission_failures_are_not_provider_failures() -> None:
"""Dispatch admission rejections must stay outside the provider vocabulary."""
workflow = _read(OPENCODE_DISPATCH_WORKFLOW)
admission_messages = (
"repository_dispatch authorization rejected",
"repository_dispatch metadata does not match the live pull request",
)
for message in admission_messages:
assert message in workflow
line = next(
candidate
for candidate in workflow.splitlines()
if message in candidate
)
assert "provider" not in line.lower()
assert "orchestrator" not in line.lower()
sidecar_index = workflow.index("Provision contextual-orchestrator review sidecar")
for message in admission_messages:
assert workflow.index(message) < sidecar_index, (
"dispatch admission runs before any sidecar exists, so its failure "
"cannot be attributed to the gateway or a provider"
)
Loading