Skip to content

feat(recipes): surface library recipes in search and run_module failures (#335) - #353

Merged
MattGyverLee merged 5 commits into
mainfrom
feat/335-surface-recipes
Oct 2, 2026
Merged

MattGyverLee merged 5 commits into
mainfrom
feat/335-surface-recipes

Conversation

@MattGyverLee

Copy link
Copy Markdown
Owner

Summary

Surfaces library recipes where the model actually looks, so it reuses a shipped recipe instead of composing from API hits (#335).

  • recommended_recipe in search_by_capability: recommend_recipe(query) looks only at shipped library recipes. It returns one only when that recipe is a clear winner: a match_terms phrase hit scoring >= 8 or a total score >= 15, the score coming from task or code words (object words alone don't count), and a lead of at least 1.3x over the runner-up. The key sits right after query, ahead of results, and carries no code, so the SC-005 rule (at most one code body) still holds. A write-intent verb (add, delete, merge, ...) together with a read-only top recipe gives no recommendation.
  • closest_recipes on run_module failures: the shared _attach_assistance_if_loop wrapper adds a "closest recipes: ..." next_steps line and a closest_recipes list in three cases: partial_module_structure rejects, casting_issues_detected rejects, and the 2nd consecutive failure with the same user_intent (tracked in SessionState.intent_failure_streak). project_locked, confirmation_required and server_state_error leave the streak alone. A recipe the submitted code already comes from is left out of the list. A nested legacy error dict gets the same keys.
  • recipe_hint: a non-blocking hint on executed runs when the intent clearly matches a library recipe but the code isn't taken from it. Code counts as taken from the recipe when at least 60% of the recipe's non-PARAMS lines appear in it.
  • Tool descriptions and docs: run_module's description gets a "RECIPES FIRST" paragraph, and search_by_capability's description mentions recommended_recipe. USAGE.md rows are updated, and docs/TOOL-CONTRACT.md documents the 3 new keys under tool-responses/1.0. All changes are additive.
  • Local recipe auto-save: a run that logged an ERROR-level report message is no longer remembered. The runner's success flag only means no exception escaped.

The coverage gaps named in the issue were already fixed on main (#337), so this PR adds no new recipe.

Closes #335

Pattern audit

  • Shape A (recipes nested where the model doesn't look): no sibling needs a fix. find_examples already returns recipes as a top-level list, list_recipes is recipe-first by design, and zero_result_fallback already has a recipes_hint.
  • Shape B (runner success flag treated as task success):
    • src/flextoolsmcp/extract_patterns.py:322 mine_operations_log keeps outcome=="ok" ops without checking report errors. Its output needs human review (source "mined", requires_human_review), so it is only a follow-up candidate.
    • local_recipes._migrate_once imports legacy skeleton rows without an outcome check. Those rows were only ever captured after successful runs, so no change.
    • build_effect_check_payload and the backup path in execution.py use success correctly for their purpose, so no change.

Tests

  • Full offline suite (.venv, -m "not requires_flex"): 4923 passed, 4 skipped, 119 deselected, 45 subtests passed.
  • New tests/test_issue335_surface_recipes.py: 41 passed.
  • python scripts/validate_integrity.py server: exit 0 (31 tools, USAGE.md complete, golden payload check passed).
  • tests/make_golden.py check: all fixtures up to date. No new error codes, so the code count is unchanged.

Live verification: none was run. These are offline-only changes.

Known merge interactions

These branches are being worked on in parallel: feat/exclusive-access-gate (#343), fix/350-352-351-write-gate, fix/preflight-recovery-dx.

  • handlers/execution.py: all of them touch it. This PR hooks only _attach_assistance_if_loop, handle_run_module (a ContextVar set at the top) and _finalize_run_module_response. It deliberately leaves the partial_module_structure builder alone, which fix/preflight-recovery-dx (dx: partial_module_structure rejects drive retry loops (identical resubmits, no assistance) #334) changes.
  • docs/TOOL-CONTRACT.md / goldens / error-code count: this PR adds no error codes, but its contract rows sit next to the ones other branches add. Expect textual conflicts, and regenerate goldens after the merge if another branch adds codes.
  • CHANGELOG.md Unreleased and the run_module tool description / USAGE.md: likely to need adjacent-line conflict resolution.
  • SessionState (new intent_failure_streak; reset_op_signals clears it): check against any session-state fields the gate branches add.

Open questions

  • Recommendation thresholds were calibrated on about 20 log queries. Should they be tuned against the eval corpus or QUERY_BATTERY before release?
  • Should recipe_hint appear only on failures, to keep successful responses smaller?
  • Should local-recipe capture be rejected on any report.Error, or only when every item errored? A bulk script that reports per-row errors for bad input is no longer remembered.

Do not merge.

🤖 Generated with Claude Code

MattGyverLee and others added 5 commits October 2, 2026 12:22
…_by_capability

Add recommend_recipe / closest_recipes / recipe_hint_for_run helpers in
server/recipes.py (shipped recipes only; no code body, so SC-005 holds).
search_by_capability emits a top-level recommended_recipe
{id, intent, params, requires_write, how_to_run} right after `query` when
one library recipe clearly wins the query.

Refs #335

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…y clean runs

- partial_module_structure / casting_issues_detected rejects, and the 2nd
  consecutive failure with the same user_intent, append a
  "closest recipes: ..." next_steps line plus a closest_recipes list. Done
  in the shared _attach_assistance_if_loop wrapper, so the
  partial_module_structure builder itself is untouched.
- Executed runs whose user_intent clearly matches a library recipe the
  code is not from carry a non-blocking recipe_hint.
- A run that called report.Error is no longer auto-saved as a local-*
  recipe (failed runs never were).

Refs #335

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Mention recipes up front in the run_module and search_by_capability tool
descriptions (USAGE.md kept in step), document recommended_recipe /
closest_recipes / recipe_hint in docs/TOOL-CONTRACT.md as additive keys
under tool-responses/1.0, add a CHANGELOG entry and tests.

The issue's coverage gaps are already covered on main:
ensure-morpheme-entries ships, and wordform-analyses matches "parse a
wordform and get morphological decomposition" (now also recommended).

closes #335

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…tents

Review fixes for #335:
- project_locked, confirmation_required and server_state_error no longer
  count toward the same-intent failure streak, so a confirm-then-locked
  recipe write run is not told to "start from" a recipe.
- closest_recipes leaves out the recipe the submitted code is already from.
- The nested legacy `error` object mirrors the next_steps line and
  closest_recipes, per the tool-responses/1.x mirror promise.
- recommend_recipe never recommends a read-only recipe for a query with a
  write verb (delete/merge/set/...), which also stops the matching
  recipe_hint.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@MattGyverLee
MattGyverLee merged commit bf852b8 into main Oct 2, 2026
2 checks passed
@MattGyverLee
MattGyverLee deleted the feat/335-surface-recipes branch October 2, 2026 18:02
MattGyverLee added a commit that referenced this pull request Oct 2, 2026
Brings in #354 (write-gate fail-opens), #353 (surface recipes), #355 (docs), #356 (preflight retry loops).

Conflict resolution keeps both sides: unknown_import (#305) and requires_exclusive_access sit side by side and the error-code count moves to 49 everywhere (response_models, both count tests, TOOL-CONTRACT, CHANGELOG). execution.py keeps main's #334 assistance-log ordering with the gate's requires_exclusive_access key. test_issue55 ladder test stubs the exclusive-access detector so it still exercises Rung 3 (gate covered in test_exclusive_access_gate.py). Full offline suite: 5273 passed, 4 skipped.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

dx: recipes are never used -- surface them in search_by_capability and failure next_steps

1 participant