Skip to content

Guard valuation questions asked in Chinese, Japanese or Korean - #25

Merged
zanwenfu merged 1 commit into
mainfrom
fix/cjk-valuation-questions
Sep 28, 2026
Merged

zanwenfu merged 1 commit into
mainfrom
fix/cjk-valuation-questions

Conversation

@zanwenfu

Copy link
Copy Markdown
Collaborator

Problem

"埃克森美孚值得买入还是持有?" ("is XOM worth buying or holding?") was not recognised as a valuation question, because the detection pattern was English only. The publication guard never ran.

The model's draft was published as written: "已有持仓可继续持有" ("keep holding if you own it"). There was no NOT RATED, for a company the engine declines to value. This predates #24.

Change

  • Recognition. The chat agent now recognises valuation questions in Chinese, Japanese and Korean: buy/sell/hold decisions, valuation, price targets and analysis, the same things the English pattern looks for. These questions now get the fixed statement plus the analysis section from Keep the run's analysis when the guard answers with its fixed statement #24, whose two checks already work in any language.
  • Published headlines. A broad answer written in Chinese, Japanese or Korean cannot be checked against a published headline, because the conflict patterns only read English. "报告评级为卖出" ("the report's rating is sell") passed while the report said HOLD. Such an answer now gets the headline itself plus the checked analysis.
  • Chinese first. A question asked in Chinese now reads its position in Chinese first. This is one deterministic line built from what the guard recorded, for example "XOM:未评级。不发布评级或公允价值:其现金流随大宗商品价格波动……" ("XOM: not rated. No rating or fair value is published: its cash flows follow commodity prices…"). The English statement follows.
    • The guard now records which statement it used: refused, withheld or published.
    • Each refusal reason has a Chinese version in summary_evidence.
  • Language of the analysis. The analysis turn now names the language ("Write it in Chinese."). When the instruction only said "the user's language", Chinese questions got English sections, and the language check refused them.
  • Chinese claim terms. Chinese value and target terms (目标价, 内在价值, 上行空间 …) count as claims only with a figure in the same clause, like their English equivalents. A disclaimer such as "无法据此量化…内在价值" ("intrinsic value cannot be quantified from this") had caused a whole section to be refused.

Verification

  • 1561 tests pass, and 5 mutations were each caught.
  • Production vs candidate on real questions:
    • Chinese valuation questions (XOM refused, META withheld, MSFT published, TSLA analyze): all are now guarded. Each reads its position in Chinese first and carries a Chinese analysis that passed both checks. I read each one by hand.
    • Unchanged: the Chinese news question and the English questions behave as before.
  • Valuation canary: PASS, with no values changed.

🤖 Generated with Claude Code

"埃克森美孚值得买入还是持有?" (is XOM worth buying or holding?) was not
recognised as a valuation question: the pattern was English only. The
publication guard never ran, and the draft was published as written: "已有
持仓可继续持有" (keep holding if you own it), with no NOT RATED, for a
company the engine declines to value.

- The chat agent recognises CJK valuation questions: buy/sell/hold
  decisions, valuation, price targets and analysis, the same things the
  English pattern looks for. They now get the fixed statement and the
  analysis section, whose two checks already read any language.
- A broad answer written in Chinese, Japanese or Korean cannot be checked
  against a published headline (the conflict patterns read English;
  "报告评级为卖出" passed while the report said HOLD), so it gets the
  headline itself plus the checked analysis.
- A question asked in Chinese reads its position in Chinese first: one
  deterministic line from what the guard recorded ("XOM:未评级。不发布
  评级或公允价值:其现金流随大宗商品价格波动……"), then the English
  statement. The guard records which statement it used (refused / withheld
  / published) to build it; the refusal reasons have Chinese twins in
  summary_evidence.

- The analysis turn names the language ("Write it in Chinese."): told
  "the user's language" in an English note, the model wrote Chinese
  questions' sections in English and the language check refused them.
- Chinese value and target terms (目标价, 内在价值, 上行空间 ...) count as
  claims with a figure in the same clause, like their English twins;
  "无法据此量化…内在价值" (intrinsic value cannot be quantified) had sunk a
  section. The instruction also names the model's valuation result and
  its gap to the market price as the fixed statement's to give.

Verified: 1561 tests; 5 mutations each caught; production vs candidate
on real Chinese questions (refused, withheld, published, analyze, news)
and English regressions: every Chinese valuation answer is now guarded,
reads its position in Chinese first, and carries a Chinese analysis that
passed both checks (read by hand); valuation canary unchanged (PASS).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@zanwenfu
zanwenfu merged commit 4e1c630 into main Sep 28, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant