feat(recipes): surface library recipes in search and run_module failures (#335) - #353
Merged
Merged
Conversation
…_by_capability
Add recommend_recipe / closest_recipes / recipe_hint_for_run helpers in
server/recipes.py (shipped recipes only; no code body, so SC-005 holds).
search_by_capability emits a top-level recommended_recipe
{id, intent, params, requires_write, how_to_run} right after `query` when
one library recipe clearly wins the query.
Refs #335
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…y clean runs - partial_module_structure / casting_issues_detected rejects, and the 2nd consecutive failure with the same user_intent, append a "closest recipes: ..." next_steps line plus a closest_recipes list. Done in the shared _attach_assistance_if_loop wrapper, so the partial_module_structure builder itself is untouched. - Executed runs whose user_intent clearly matches a library recipe the code is not from carry a non-blocking recipe_hint. - A run that called report.Error is no longer auto-saved as a local-* recipe (failed runs never were). Refs #335 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Mention recipes up front in the run_module and search_by_capability tool descriptions (USAGE.md kept in step), document recommended_recipe / closest_recipes / recipe_hint in docs/TOOL-CONTRACT.md as additive keys under tool-responses/1.0, add a CHANGELOG entry and tests. The issue's coverage gaps are already covered on main: ensure-morpheme-entries ships, and wordform-analyses matches "parse a wordform and get morphological decomposition" (now also recommended). closes #335 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…tents Review fixes for #335: - project_locked, confirmation_required and server_state_error no longer count toward the same-intent failure streak, so a confirm-then-locked recipe write run is not told to "start from" a recipe. - closest_recipes leaves out the recipe the submitted code is already from. - The nested legacy `error` object mirrors the next_steps line and closest_recipes, per the tool-responses/1.x mirror promise. - recommend_recipe never recommends a read-only recipe for a query with a write verb (delete/merge/set/...), which also stops the matching recipe_hint. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
MattGyverLee
added a commit
that referenced
this pull request
Oct 2, 2026
Brings in #354 (write-gate fail-opens), #353 (surface recipes), #355 (docs), #356 (preflight retry loops). Conflict resolution keeps both sides: unknown_import (#305) and requires_exclusive_access sit side by side and the error-code count moves to 49 everywhere (response_models, both count tests, TOOL-CONTRACT, CHANGELOG). execution.py keeps main's #334 assistance-log ordering with the gate's requires_exclusive_access key. test_issue55 ladder test stubs the exclusive-access detector so it still exercises Rung 3 (gate covered in test_exclusive_access_gate.py). Full offline suite: 5273 passed, 4 skipped.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Surfaces library recipes where the model actually looks, so it reuses a shipped recipe instead of composing from API hits (#335).
recommended_recipein search_by_capability:recommend_recipe(query)looks only at shipped library recipes. It returns one only when that recipe is a clear winner: a match_terms phrase hit scoring >= 8 or a total score >= 15, the score coming from task or code words (object words alone don't count), and a lead of at least 1.3x over the runner-up. The key sits right afterquery, ahead ofresults, and carries nocode, so the SC-005 rule (at most one code body) still holds. A write-intent verb (add, delete, merge, ...) together with a read-only top recipe gives no recommendation.closest_recipeson run_module failures: the shared_attach_assistance_if_loopwrapper adds a "closest recipes: ..." next_steps line and aclosest_recipeslist in three cases: partial_module_structure rejects, casting_issues_detected rejects, and the 2nd consecutive failure with the same user_intent (tracked inSessionState.intent_failure_streak).project_locked,confirmation_requiredandserver_state_errorleave the streak alone. A recipe the submitted code already comes from is left out of the list. A nested legacyerrordict gets the same keys.recipe_hint: a non-blocking hint on executed runs when the intent clearly matches a library recipe but the code isn't taken from it. Code counts as taken from the recipe when at least 60% of the recipe's non-PARAMS lines appear in it.The coverage gaps named in the issue were already fixed on main (#337), so this PR adds no new recipe.
Closes #335
Pattern audit
src/flextoolsmcp/extract_patterns.py:322mine_operations_logkeeps outcome=="ok" ops without checking report errors. Its output needs human review (source "mined", requires_human_review), so it is only a follow-up candidate.local_recipes._migrate_onceimports legacy skeleton rows without an outcome check. Those rows were only ever captured after successful runs, so no change.build_effect_check_payloadand the backup path in execution.py use success correctly for their purpose, so no change.Tests
.venv,-m "not requires_flex"): 4923 passed, 4 skipped, 119 deselected, 45 subtests passed.tests/test_issue335_surface_recipes.py: 41 passed.python scripts/validate_integrity.py server: exit 0 (31 tools, USAGE.md complete, golden payload check passed).tests/make_golden.py check: all fixtures up to date. No new error codes, so the code count is unchanged.Live verification: none was run. These are offline-only changes.
Known merge interactions
These branches are being worked on in parallel: feat/exclusive-access-gate (#343), fix/350-352-351-write-gate, fix/preflight-recovery-dx.
handlers/execution.py: all of them touch it. This PR hooks only_attach_assistance_if_loop,handle_run_module(a ContextVar set at the top) and_finalize_run_module_response. It deliberately leaves the partial_module_structure builder alone, which fix/preflight-recovery-dx (dx: partial_module_structure rejects drive retry loops (identical resubmits, no assistance) #334) changes.docs/TOOL-CONTRACT.md/ goldens / error-code count: this PR adds no error codes, but its contract rows sit next to the ones other branches add. Expect textual conflicts, and regenerate goldens after the merge if another branch adds codes.CHANGELOG.mdUnreleased and therun_moduletool description / USAGE.md: likely to need adjacent-line conflict resolution.SessionState(newintent_failure_streak;reset_op_signalsclears it): check against any session-state fields the gate branches add.Open questions
Do not merge.
🤖 Generated with Claude Code