Skip to content

Give VYNN's answer with a confidence alert instead of "NOT RATED" - #30

Merged
zanwenfu merged 1 commit into
mainfrom
feat/confidence-alerts
Oct 2, 2026
Merged

zanwenfu merged 1 commit into
mainfrom
feat/confidence-alerts

Conversation

@zanwenfu

@zanwenfu zanwenfu commented Oct 2, 2026

Copy link
Copy Markdown
Collaborator

What this does

A sound model that well-covered analysts did not back was withheld, and the user read "NOT RATED". A valuation has no certain answer, only better or worse evidence. The engine now gives its own answer and says plainly how far it stands from the Street:

Low confidence: VYNN's fair value is 49% below the market price, while the mean target of 35 analysts is 15% above it. That is a large gap between VYNN and the Street, so treat this as VYNN's own view and weigh both.

The model is never pulled toward consensus, and consensus is never hidden. Analyst targets and ratings still never enter intrinsic-value arithmetic.

The outcomes

Outcome When What is published
Published The Street corroborates the model. Fair value, rating, target. Unchanged.
Flagged (was "NOT RATED") The model is sound and only the Street stands apart: its target is on the other side of the price, sits at the price, backs under half of the move, its rating points another way, or no current target exists to check a DCF-only value against. Fair value, rating and target at low confidence, with a confidence alert that states both positions.
No single fair value (was "NOT RATED") The model itself supports none: a failed method, methods more than 1.8x apart, no positive value, stale statements, a method the company does not fit, a near-term forecast or margin far from analysts' estimates. What the methods do support, named for what it is, and the reason.
Evidence only (was "NOT RATED") No cash flows to value: commodity producers, REITs, insurers, funds, crypto. Same answer as before, without the label.

"NOT RATED" survives as the machine value of rating only. No surface prints it, and the answer guard rejects "not rated", "unrated" and "withheld" in generated prose.

Words follow what the methods support

One estimate is never called a range (src/confidence_alert.py, support_shape):

Methods support Report heading Chat
a range ### Investment View: Range Only "a range, not a single fair value, because ..."
one estimate ### Investment View: Scenario Estimate Only "a scenario estimate, not a fair value, because ..."
nothing usable ### Investment View: No Fair Value "VYNN has no fair value to state here, because ..."
method ruled out (range or estimate, as above) "a scenario, not a fair value, because ..."

A flagged report keeps ### Investment Rating: X and adds **Rating Confidence**: Low and **Confidence Alert**: ... on the first page, in the valuation section and under the rating. The workbook's status cell reads PUBLISHABLE (low confidence: far from analyst consensus) with the alert sentence in the reason cell. _vynn.valuation_publication.confidence_alert carries the alert with the sentence and label every surface prints; the schema version is unchanged.

Guardrails

  • An alert qualifies a published value only, at every layer.
  • If an alert cannot be built for a Street-only disagreement, the corporate and bank boundaries fall back to no single value, never to an unflagged one.
  • When a later check blocks the run, the analysts' position follows the reason as context and is never the reason.
  • If the generated answer leaves the alert out, the code adds it, in English or Chinese to match the answer.
  • The artifact audit fails a workbook that publishes a flagged value without saying so.

Release gate

  • The canary summary reports PUBLISHED, FLAGGED and WITHHELD.
  • scripts/valuation_canary_expectations.json predicts META, TSLA, AMD, AAPL, AMZN and PYPL as FLAGGED. That is a prediction from the last basket run. It must be confirmed by running the candidate image against the deployed one before this is deployed.
  • The publish-rate floor rises from 0.30 to 0.55. Install the nightly script together with this image; against the current image, which publishes 5 of 16, the new floor fails.

Checked

  • 1,911 tests (1,729 before) on Python 3.11 and 3.12.
  • The deployed decision code and this one, side by side on 20,000 random inputs for each of the corporate boundary, the bank boundary and the rating calculator: published and model-blocked cases are identical, and every changed case is a Street-only disagreement that is now published with an alert.
  • Six workbooks built offline from saved statements pass the artifact audit and read back through the API's summary reader with the expected rating, range and alert.

Before merging: three calls

  1. Very large gaps publish too. The rule has no upper limit. A cash-flow value far below the price (Tesla is the clearest case) now reads as a sell at low confidence with the alert, where it read "NOT RATED".
  2. Uncovered names publish as "not yet confirmed by analysts" when a DCF-only value is 15% or more from the price.
  3. A near-term revenue forecast more than 15% from analysts' estimates still gives no single value, because that usually means an input is off and not a difference of opinion. It can be flagged instead if preferred.

Deploy order

The API, then the web app, then this image. Both already read the old and the new wording.

🤖 Generated with Claude Code

https://claude.ai/code/session_01PnXmSYpqKrQVxdzGZyq9Ct

A sound model that well-covered analysts did not back was withheld, and the
user read "NOT RATED". A valuation has no certain answer, only better or worse
evidence, so the engine now publishes its own fair value and rating and states
how far it stands from the Street: "Low confidence: VYNN's fair value is 49%
below the market price, while the mean target of 35 analysts is 15% above it.
That is a large gap between VYNN and the Street, so treat this as VYNN's own
view and weigh both." The model is never pulled toward consensus, and
consensus is never hidden. Analyst targets and ratings still never enter
intrinsic-value arithmetic.

What decides the outcome (src/confidence_alert.py, the publication boundary):

- published: the Street corroborates the model. Unchanged.
- flagged: the model is sound and only the Street stands apart. Published,
  rating confidence low, with a confidence alert that states both positions.
  The alert is in the chat answer, on the report's first page, in its
  valuation section and under its rating, in the workbook's status and reason
  cells, in the publication metadata the API reads, and in the findings line.
- no single fair value: the model itself supports none (a failed method,
  methods more than 1.8x apart, no positive value, stale statements, a method
  the company does not fit, a near-term forecast or margin far from analysts'
  estimates). Unchanged in substance. The words now say what the answer is,
  and follow what the methods support, so one estimate is never called a
  range: "a range, not a single fair value", "a scenario estimate, not a fair
  value", "a scenario, not a fair value", or "no fair value to state".
- evidence only (commodity producers, REITs, insurers, funds, crypto):
  unchanged, without the label.

"NOT RATED" survives as the machine value of `rating` only. No surface prints
it, and the answer guard rejects "not rated", "unrated" and "withheld" in
generated prose.

Report wording the API and the web parse (both read the old wording too):
  ### Investment Rating: X        or  ### Investment View: Range Only |
                                      Scenario Estimate Only | No Fair Value |
                                      No Market Price
  **Single Fair Value**: Range only | Scenario estimate only | None
  **Supported Valuation Range**: $a – $b   or   **Supported Scenario Estimate**: $a
  **Rating Confidence**: Low  +  **Confidence Alert**: <both positions>
Workbook status cell: PUBLISHABLE | PUBLISHABLE (low confidence: ...) |
SCENARIO RANGE ONLY | SCENARIO ESTIMATE ONLY.
`_vynn.valuation_publication.confidence_alert` carries the alert with the
sentence and label every surface prints; the schema version is unchanged.

Guardrails:
- an alert qualifies a published value only, at every layer;
- if an alert cannot be built for a Street-only disagreement, the corporate
  and bank boundaries fall back to no single value, never to an unflagged one;
- when a later check blocks the run, the analysts' position follows the reason
  as context and is never the reason;
- the artifact audit fails a workbook that publishes a flagged value without
  saying so.

Release gate: the canary summary reports PUBLISHED, FLAGGED and WITHHELD.
scripts/valuation_canary_expectations.json predicts META, TSLA, AMD, AAPL,
AMZN and PYPL as FLAGGED; that prediction must be confirmed on the server
before this image is deployed. The publish-rate floor rises from 0.30 to 0.55,
so the nightly script must be installed together with this image.

Checked: 1,911 tests (1,729 before) on Python 3.11 and 3.12. The deployed
decision code and this one were run side by side on 20,000 random inputs for
each of the corporate boundary, the bank boundary and the rating calculator:
published and model-blocked cases are identical, and every changed case is a
Street-only disagreement that is now published with an alert. Six workbooks
built offline from saved statements pass the artifact audit and read back
through the API's summary reader.

Deploy order: api-runner, then the web app, then this image.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PnXmSYpqKrQVxdzGZyq9Ct
Copilot AI balanced review requested due to automatic review settings October 2, 2026 06:48

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@zanwenfu
zanwenfu merged commit 050eeca into main Oct 2, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants