Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 6 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -164,6 +164,8 @@ The Excel model is the same idea made tangible: **all formulas are live, not sta

The harder discipline is that **a number the engine computes correctly can still be meaningless.** A fair value averaged from methods that contradict each other is arithmetically valid and analytically worthless, and it is the most dangerous output the system can produce, because it looks exactly like a precise answer. Two rails address this: the valuation legs are made to *converge by construction* (see [The DCF engine](#the-dcf-engine)), and their remaining spread is classified and reported. When the methods disagree the answer leads with a range; when one fails outright, it says so instead of quietly presenting the survivor as a consensus.

A third case is different in kind: **the model is sound, and well-covered analysts do not back its conclusion.** A valuation has no certain answer, only better or worse evidence, so the engine does not hide its own answer behind the Street's, and does not pull it toward consensus either. It publishes the fair value and rating with a **confidence alert** that states both positions side by side ("VYNN's fair value is 49% below the market price, while the mean target of 35 analysts is 15% above it"), sets the rating confidence to low, and prints the same alert on the report's first page, in the workbook's status cell and in the chat answer (`src/confidence_alert.py`). Analyst targets and ratings remain a benchmark: they never enter intrinsic-value arithmetic.

### Valuation calibration benchmark

Every release is checked against a fixed basket before the worker image is
Expand All @@ -177,7 +179,9 @@ the Street target and the gap to market per name, and fails when the publish
rate falls below threshold or any name errors. Between 2026-09-15 and
2026-09-26 the engine withheld 9 of 11 production runs and nothing caught it;
this check exists so a publish-rate collapse fails a build instead of reaching
a user. Before a deploy the same summary runs with `--expect
a user. The summary reports three outcomes per name: PUBLISHED (the Street
corroborates the model), FLAGGED (published with a confidence alert) and
WITHHELD (the gate's word for a model that supports only a range). Before a deploy the same summary runs with `--expect
scripts/valuation_canary_expectations.json`, which names the outcome every
basket name is supposed to have and why; a candidate ships with zero
unexplained differences, or the expectation changes in the same commit as the
Expand Down Expand Up @@ -544,7 +548,7 @@ prompts/ # 34 externalized prompt templates
## Known limitations

- **News freshness.** SerpAPI's Google News results can lag breaking news by 15–30 minutes; not suitable for intraday signals.
- **Model calibration is not yet an accuracy claim.** Established-company inputs are now grounded and exceptional/uncorroborated outputs are withheld, but the rating weights and difficult profiles (pre-revenue biotech, SPACs, recent IPOs with thin history) still require a clean, versioned cross-sectional cohort and a genuine 12-month outcome backtest before they can be called calibrated.
- **Model calibration is not yet an accuracy claim.** Established-company inputs are now grounded, an output the model itself cannot support is shown as a range, and an output well-covered analysts do not back is published with a confidence alert, but the rating weights and difficult profiles (pre-revenue biotech, SPACs, recent IPOs with thin history) still require a clean, versioned cross-sectional cohort and a genuine 12-month outcome backtest before they can be called calibrated.
- **Companies a DCF does not fit.** Pre-revenue and deeply FCF-negative businesses yield negative intrinsic values under both DCF methods; no assumption set repairs this, because discounted cash flow is the wrong instrument for them. The blend excludes failed legs and the dispersion rail states plainly when a fair value rests on one surviving method — but the honest output in these cases is a range and a caveat, not a price target.
- **Yahoo Finance rate limiting.** `yfinance` can throttle under heavy concurrent use; the client retries with backoff but does not queue requests across simultaneous analyses.
- **Symbol resolution.** Non-Latin names are resolved via the model's transliteration plus search; obscure or ambiguously-named companies may need the ticker stated explicitly.
Expand Down
3 changes: 2 additions & 1 deletion main.py
Original file line number Diff line number Diff line change
Expand Up @@ -415,7 +415,8 @@ def run_model_generation_stage(self) -> Dict:
f"(P/B {bank_inputs.get('justified_pb', 0):.2f}; single method)"
)
publication_note = (
"point estimate withheld" if valuation_override.get("point_estimate_withheld")
"no single fair value"
if valuation_override.get("point_estimate_withheld")
else "point estimate publishable"
)
self.logger.info(
Expand Down
5 changes: 3 additions & 2 deletions prompts/answer_synthesis.md
Original file line number Diff line number Diff line change
Expand Up @@ -23,10 +23,11 @@ Write a direct answer to the user's question. This is the message they read —
- **Match the length to the question.** A narrow question ("what's the P/E?", "how's sentiment?") deserves 1–3 sentences. A broad request ("analyze this stock", "should I invest") deserves a fuller 5–8 sentence answer covering valuation, catalysts, risks, and a recommendation.
- **Write like you're talking to the user**, in second person where natural ("Your main concern here should be…"). Warm, precise, senior-analyst voice. No preamble like "Based on my analysis" or "I have completed" — just answer.
- **Be balanced and honest.** Note both the bullish and bearish side when it's relevant to the question. If the data is thin or mixed, say that rather than overclaiming.
- If the valuation summary says the point estimate/rating was withheld, do not quote the internal midpoint or an upside/downside to it. State the supported method range and the publication reason. If reverse-DCF evidence is supplied, explain what today's price requires from future cash flow.
- If the valuation summary says there is no single fair value for this run, do not quote the internal midpoint or an upside/downside to it. Say what VYNN's answer is instead, in these words: "a range, not a single fair value" when a method range is listed, or "a scenario estimate, not a fair value" when a single scenario estimate is listed. State those figures and the reason. Never call one estimate a range, and never call the result not rated, unrated or withheld. If reverse-DCF evidence is supplied, explain what today's price requires from future cash flow.
- If the valuation summary carries a confidence alert, the fair value and rating stand as VYNN's own answer at low confidence. State both positions from the alert (VYNN's gap to the market and the analysts') early in the answer, and never describe the result as confirmed by, or consistent with, analyst consensus.
- Treat the independent human-analyst benchmark as a required cross-check when it is supplied: compare the model with the consensus target and FY1/FY2 revenue estimates, and explain whether they agree. Never blend an analyst price target into intrinsic value or present consensus as proof that the market is right.
- When current licensed analyst record metadata is supplied, attribute the available firm/date/action/rating/target records and use them as external context; do not merely say analysts were considered. Licensed rationale prose was not read or collected, so never invent rationale themes or imply that metadata reveals an analyst's full reasoning.
- Call article-derived tone **news sentiment**, never "overall sentiment" or an investment rating. A bearish news screen is not a SELL, and bullish analyst consensus is not a BUY. If the model is NOT RATED, do not manufacture a directional recommendation from either one.
- Call article-derived tone **news sentiment**, never "overall sentiment" or an investment rating. A bearish news screen is not a SELL, and bullish analyst consensus is not a BUY. If the model states no rating, do not manufacture a directional recommendation from either one.
- Never describe sentiment as bullish or bearish when freshness coverage is limited or unavailable.
- The full model (.xlsx) and report (.md) are attached separately for them, so you don't need to tell them "a report was generated" — reference the *findings*, not the artifacts.

Expand Down
13 changes: 7 additions & 6 deletions prompts/supervisor_performance_summary.md
Original file line number Diff line number Diff line change
Expand Up @@ -29,16 +29,17 @@ Write a comprehensive 5-7 sentence summary that:
6. **Provides clear recommendation** - Buy, hold, sell, or what analysis is still needed?

**Critical Rules:**
- If the valuation evidence says a point estimate or rating was withheld, do not quote the internal midpoint, upside/downside to it, or invent BUY/HOLD/SELL. Give the supported method range and the exact reason.
- If the valuation evidence says there is no single fair value for this run, do not quote the internal midpoint, upside/downside to it, or invent BUY/HOLD/SELL. Say what VYNN's answer is instead, in these words: "a range, not a single fair value" when a method range is listed, or "a scenario estimate, not a fair value" when a single scenario estimate is listed. Give those figures and the exact reason. Never call one estimate a range, and never call the result not rated, unrated or withheld.
- If the valuation evidence carries a confidence alert, the fair value and rating stand as VYNN's own answer at low confidence: state both positions from the alert (VYNN's gap to the market and the analysts'), and never describe the result as confirmed by analyst consensus.
- Use a supplied human-analyst benchmark as an explicit cross-check, including model-versus-Street revenue gaps. It is external evidence, never an intrinsic-value input.
- Call article-derived tone "news sentiment," never "overall sentiment" or an investment rating. Do not turn news tone or analyst consensus into a recommendation when the model is NOT RATED.
- Call article-derived tone "news sentiment," never "overall sentiment" or an investment rating. Do not turn news tone or analyst consensus into a recommendation when the model states no rating.
- Start with the finding, NOT the process ("Meta's DCF shows...", not "We generated a DCF...")
- Use SPECIFIC NUMBERS from the data above. For a publishable model, cite fair
value, current price, implied return, WACC, terminal growth, and revenue
growth. For a withheld model, cite only the supported method range, current
price, assumptions, and publication reason—never the internal midpoint or
implied return.
- **Financial Model**: Publishable value/return or withheld range/reason, Current Price, WACC, Terminal Growth, Revenue Growth rates
growth. For a model with no single fair value, cite only the supported method
range or scenario estimate, current price, assumptions, and the reason, never
the internal midpoint or implied return.
- **Financial Model**: Publishable value/return (with its confidence alert, if any) or the supported range or scenario estimate and its reason, Current Price, WACC, Terminal Growth, Revenue Growth rates
- **News Analysis**: Number of articles, Sentiment, Specific catalyst descriptions, Specific risk descriptions, Severity/Timeline/Impact details
- Never use placeholders like X, Y, N/A - if data is missing, acknowledge it directly
- Be comprehensive but focused (5-7 sentences, not more)
Expand Down
13 changes: 8 additions & 5 deletions scripts/nightly_valuation_canary.sh
Original file line number Diff line number Diff line change
Expand Up @@ -21,11 +21,14 @@ set -euo pipefail
DEPLOY_DIR=${DEPLOY_DIR:-/opt/vynn/deploy}
OUT_ROOT=${OUT_ROOT:-/var/lib/vynn/canary}
BASKET=${BASKET:-"TSLA AMD NVDA META AAPL AMZN GOOGL MSFT CRH MC.PA PYPL PCJEWELLER.NS MU GM TEX BKNG"}
# Ten basket names are withheld by design and Booking sits on a boundary
# (scripts/valuation_canary_expectations.json), so 5 of 16 publish on a
# healthy engine, 6 on some days. The floor fails below 5 of 16: it catches a
# collapse, not one name flipping, which the pre-deploy --expect gate catches.
MIN_PUBLISH_RATE=${MIN_PUBLISH_RATE:-0.30}
# Four basket names are range-only by design and Booking sits on a boundary
# (scripts/valuation_canary_expectations.json), so 11 of 16 publish on a
# healthy engine (six of them with a confidence alert), 12 on some days. The
# floor fails below 9 of 16: it catches a collapse, not one name flipping,
# which the pre-deploy --expect gate catches. Install this file together with
# the engine image that publishes flagged names: against an older image, which
# withheld them, 5 of 16 publish and this floor fails.
MIN_PUBLISH_RATE=${MIN_PUBLISH_RATE:-0.55}
# Every basket name is an operating company, so any refusal is a regression.
MAX_REFUSED=${MAX_REFUSED:-0}
# Nightly output is about 10 MB; keep a month of it.
Expand Down
37 changes: 19 additions & 18 deletions scripts/valuation_canary_expectations.json
Original file line number Diff line number Diff line change
@@ -1,5 +1,5 @@
{
"_comment": "Expected publication outcome per basket name for the pre-deploy gate (scripts/valuation_canary_summary.py --expect). Change an entry only together with the engine change that explains it, and say why here. 2026-09-27: calibrated premium (3.24%) is the default.",
"_comment": "Expected publication outcome per basket name for the pre-deploy gate (scripts/valuation_canary_summary.py --expect). Change an entry only together with the engine change that explains it, and say why here. 2026-09-27: calibrated premium (3.24%) is the default. 2026-10-02: a sound model that well-covered analysts do not back is published with a confidence alert (FLAGGED) instead of withheld; WITHHELD now means the model itself supports no single value (a failed or contradicting method, stale statements, a method the company does not fit) and the answer is the range.",
"NVDA": {
"status": "PUBLISHED",
"why": "Calibrated premium lifts the DCF to about +57%; the 59-analyst Street target (+46%) covers more than half the move in the same direction."
Expand All @@ -9,8 +9,8 @@
"why": "DCF about +14% with the calibrated premium; the Street (+12%) corroborates if it lands above 15%."
},
"META": {
"status": "WITHHELD",
"why": "Calibrated premium lifts the DCF to about +23%; the Street target (+5%) covers less than half the move, so the boundary withholds and the answer leads with the model view."
"status": "FLAGGED",
"why": "Calibrated premium lifts the DCF to about +23%; the Street target (+5%) covers less than half the move. Published with a confidence alert since 2026-10-02 (was withheld)."
},
"GOOGL": {
"status": "PUBLISHED",
Expand All @@ -21,40 +21,40 @@
"why": "DCF about +25% with the calibrated premium; the Street (+57%) corroborates."
},
"TSLA": {
"status": "WITHHELD",
"why": "Consensus cash flows support a value far below the market; exit leg unavailable above the 80x boundary. Model view is shown, no rating."
"status": "FLAGGED",
"why": "Consensus cash flows support a value far below the market while the Street target sits above it; the exit leg is unavailable above the 80x boundary, which omits it rather than failing it. Published with a confidence alert since 2026-10-02 (was withheld)."
},
"AMD": {
"status": "WITHHELD",
"why": "DCF far below market on a 30%-capped fade; the Street target sits at the market, so nothing corroborates the gap."
"status": "FLAGGED",
"why": "DCF far below market on a 30%-capped fade; the Street target sits at the market, so nothing corroborates the gap. Published with a confidence alert since 2026-10-02 (was withheld)."
},
"AAPL": {
"status": "WITHHELD",
"why": "DCF-only, about -38% vs market; Street at the market."
"status": "FLAGGED",
"why": "DCF-only, about -38% vs market; Street at the market. Published with a confidence alert since 2026-10-02 (was withheld)."
},
"AMZN": {
"status": "WITHHELD",
"why": "DCF about -34% vs market while the Street is +32%; model and Street disagree."
"status": "FLAGGED",
"why": "DCF about -34% vs market while the Street is +32%; model and Street point opposite ways. Published with a confidence alert since 2026-10-02 (was withheld)."
},
"PYPL": {
"status": "WITHHELD",
"why": "DCF well above market while the Street target sits at the market."
"status": "FLAGGED",
"why": "DCF well above market while the Street target sits at the market. Published with a confidence alert since 2026-10-02 (was withheld)."
},
"MC.PA": {
"status": "WITHHELD",
"why": "Latest annual filing beyond the 200-day limit and no semiannual bridge."
"why": "Latest annual filing beyond the 200-day limit and no semiannual bridge. Range only: stale statements are a model input problem, not a disagreement with the Street."
},
"PCJEWELLER.NS": {
"status": "WITHHELD",
"why": "Perpetual DCF is negative; a failed leg never publishes."
"why": "Perpetual DCF is negative; a failed leg never publishes. Range only."
},
"MU": {
"status": "WITHHELD",
"why": "Memory maker: DRAM/NAND prices set the covered years' margin, so the model is a scenario until a mid-cycle margin exists."
"why": "Memory maker: DRAM/NAND prices set the covered years' margin, so the model is a scenario until a mid-cycle margin exists. Range only."
},
"GM": {
"status": "WITHHELD",
"why": "GM Financial is a consolidated captive lender; scenario only until an operating/finance split exists."
"why": "GM Financial is a consolidated captive lender; scenario only until an operating/finance split exists. Range only."
},
"TEX": {
"status": "PUBLISHED",
Expand All @@ -63,8 +63,9 @@
"BKNG": {
"status": [
"PUBLISHED",
"FLAGGED",
"WITHHELD"
],
"why": "Reports no cost of revenue in any year (expenses by function). It must build and pass integrity: before 2026-09-27 every Booking model failed the historical-period check, and payables measured against the absent line were projected at zero. Either publication outcome is accepted because the methods span right at the 1.8x limit: DCF about +34%, comps about +134% on Airbnb, Hilton and Royal Caribbean at 26.7x EBITDA versus Booking's own 12.2x."
"why": "Reports no cost of revenue in any year (expenses by function). It must build and pass integrity: before 2026-09-27 every Booking model failed the historical-period check, and payables measured against the absent line were projected at zero. Any publication outcome is accepted because the methods span right at the 1.8x limit: DCF about +34%, comps about +134% on Airbnb, Hilton and Royal Caribbean at 26.7x EBITDA versus Booking's own 12.2x. When the methods land inside the limit the value publishes, flagged if the Street does not back it."
}
}
Loading
Loading