diff --git a/.github/agents/content-scout.agent.md b/.github/agents/content-scout.agent.md index 3c6e669..a04bb4f 100644 --- a/.github/agents/content-scout.agent.md +++ b/.github/agents/content-scout.agent.md @@ -1245,6 +1245,23 @@ Where the platform exposes them, also include engagement metrics (likes/upvotes/ |------------|---------------|-----------|-------------------|---------------| + + ## Launch Coverage Tracker diff --git a/.github/prompts/scout-scan.prompt.md b/.github/prompts/scout-scan.prompt.md index 4889fa4..1b549f2 100644 --- a/.github/prompts/scout-scan.prompt.md +++ b/.github/prompts/scout-scan.prompt.md @@ -96,6 +96,7 @@ Run a content scan using the Content Scout agent. - If no fresh sidecars exist, attempt a refresh: run `node tools/browser-scan/index.mjs scan --slug {slug}` for each subject. If the command exits with a "no browser on CDP port" error, show the tip below **once** and continue β€” the API/RSS layers still cover X, LinkedIn, and Reddit with adequate fidelity. - One-time setup tip (show only when CDP is not running): *"For higher-fidelity X / LinkedIn / Reddit coverage, run `node tools/browser-scan/launch-edge.mjs` once, sign in to all three platforms, and leave the browser open. Future `/scout-scan` runs will use it automatically."* - When sidecars are available, ingest the freshest per platform and tag those items `*-browser` provenance β€” they take priority over Brave/RSS/old.reddit results, deduped by permalink. This includes the `*-content-sites.json` sidecar, which covers **Microsoft Tech Community, DZone, C# Corner, and Hashnode** (sources the API/RSS layers can't reach due to login wall / anti-bot / 500 / 404); its items carry `subSource` of `techcommunity` / `dzone` / `csharpcorner` / `hashnode` and feed the **content sections** (never Mindshare). + - When `Competitor tracking` is `on` and the config has a `## Competitors` section, the browser-scan also emits a `{stamp}-competitors.json` sidecar β€” deduplicated Reddit + X items that name a tracked competitor. Each item includes canonical `competitor` / `competitorMatches`, target-scoped `competitorSentiment` / `competitorSentimentConfidence`, `platform`, `timestamp`, `url`, and a separate `switchingDirection`. Feed these into the `## Competitor & Market Signals` section only (never the content sections or Mindshare). The `{stamp}-meta.json` sidecar's `competitors` block carries isolated per-competitor and per-source mention/sentiment aggregates plus non-fatal `sourceFailures`. - Search all enabled networks using the configured search terms. **Reddit, X/Twitter, LinkedIn, and Bluesky are always attempted** β€” never skipped just because one credential is missing (see "API Keys" exceptions in the agent definition for which layer/fallback to use when a key is empty). For Bluesky specifically: if both `BLUESKY_HANDLE` and `BLUESKY_APP_PASSWORD` are set in `.env`, you MUST call `createSession` and run the search; never write "credentials present but API call not completed." - **Open-web blog discovery is mandatory, not just RSS-by-tag.** The named blog platforms (Dev.to, Medium, Hashnode, …) miss self-hosted and vendor blogs. Run an **unrestricted** Brave web search (`q={term}` with no `site:` filter, `freshness=pm`) **once for every term in the config's `### Search Terms (text)` list** (not just one product term), drop hosts already covered by a dedicated layer (reddit/x/linkedin/youtube/github/stackoverflow/dev.to/medium/hashnode), and treat the rest as candidate blog posts to date-gate, relevancy-check, and score. Additionally, for any `## Known Author Watchlist` / `## Influencers to Monitor` entry that lists a blog domain (e.g., `benday.com/blog`), run a targeted `site:{domain}` Brave query per search term so that author's posts are caught. See "Blog Sources β†’ Open-web blog discovery" in the agent definition. - Apply the content quality filter (date gate + relevancy gate). **Drop all hiring/recruiting/job-search content from EVERY section** (numbered tables, Conversations, Feature Requests, Influence Movers, social posts) per the "No Hiring Content" hard rule in the agent doc β€” even when the post mentions the product, even when the author is on the known-author list. Group these drops under a single `hiring/recruiting` counter in the JSON sidecar's `drop_reasons`; do not enumerate them in the user-facing summary. @@ -110,6 +111,7 @@ Run a content scan using the Content Scout agent. - **Open-CFP gate (mandatory).** A conference belongs in `## Open Calls for Papers (CFPs)` **only if you have verified its CFP is still open** β€” i.e. it has a real submission page (Sessionize / Pretalx / typeform with a review process, per the CFP Vetting Checklist) AND a close date that is today or later. **Fetch the CFP page and confirm the close date before listing it** β€” a stale "CFP now open" banner does not count (e.g., a page may say "open" while its own milestones show the close date already passed). If you cannot verify an open close date β‰₯ today, the CFP does not go in the report. Never list a blog title, talk recap, or past event as a CFP. - When you confirm a close date during a scan, write it back to the config's `CFP Closes` column so future scans inherit it. - Populate the `## Mindshare` section with **date-gated** community social listening for THIS report's window only: include a post only if its publish date falls inside the report period. Never carry forward older posts to pad it. Conversation platforms only (Bluesky / X / LinkedIn / Reddit); exclude blogs, YouTube, GitHub, Stack Overflow, and official/owned handles. Do not add scan-date or rolling-window caveats to the body. The **X / LinkedIn / Reddit rows come primarily from the Layer 0 browser-scan sidecars** ingested in the browser-scan step above (`*-x.json` / `*-linkedin.json` / `*-reddit.json`); Bluesky rows come from the `searchPosts` API. Browser-scan `google-news` / `google-web` sidecar items are NOT social listening β€” they feed the content sections, never Mindshare. + - If **Competitor tracking** is enabled in config (`- **Competitor tracking:** on`) and the config's `## Competitors` section lists one or more products, populate the report's `## Competitor & Market Signals` section for those products (match on each competitor's name **and** its listed aliases). **Reddit + X competitor mentions come from the Layer 0 browser-scan `{stamp}-competitors.json` sidecar** produced in Step 0 β€” ingest it first. Then cover the platforms browser-scan doesn't reach by running **per-alias API queries** on Hacker News (Algolia `search_by_date`), Stack Overflow, and Bluesky (`searchPosts`), **plus each competitor's own blog / release-notes / changelog** for the window. Canonicalize and deduplicate every URL/permalink before classification. Classify the author's stance toward each matched competitor independently from stance toward our product; never copy the primary-product sentiment onto a competitor. Store switching separately as `competitor_to_primary` (our win), `primary_to_competitor` (our loss), `competitor_to_competitor` (neutral market movement unless the text explicitly evaluates the matched competitor), or `none`. A failed competitor source is additive and non-fatal: record it and continue the primary scan and remaining competitor sources. In the report JSON sidecar, persist `competitor_mentions` with competitor identity/matches, competitor sentiment/confidence, platform, timestamp, URL, and switching direction; persist isolated `competitor_aggregates` by competitor and source and `competitor_source_failures`. Never add competitor sentiment counts to primary `sentiment_totals`. Fill one table row per competitor with β€” **Content Volume** (rough count of in-window mentions), **Sentiment** (community stance *toward that competitor*: 🟒 favorable / βšͺ mixed / πŸ”΄ unfavorable), **Switching Signals** (migration posts to/from the competitor, scored from OUR product's perspective per the Directional rule: FROM competitor β†’ us = 🟒 win, FROM us β†’ competitor = πŸ”΄ loss), and **Notable Items** (1–3 announcements, GA launches, outages, pricing changes) with **validated** links. Keep this **inline in the one content report** β€” do NOT write a separate `-competitors.md` file during a scan (see "One Scan = One Report"). A dedicated deep-dive competitor report is a separate, on-demand artifact. The web UI's **Competitors** tab surfaces this section automatically. - **Validate every URL before it goes in the report (mandatory).** Never present a link you have not confirmed is real, reachable content. Run `node tools/validate-urls.mjs ` (or call the same logic via `tools/lib/url-validate.mjs`) over every external URL you intend to include. Drop or replace any link the check marks `DEAD` (HTTP 404/410, malformed, or a known-broken shape such as a LinkedIn `/feed/sdui-post/` permalink). Prefer the canonical source over an aggregator/redirect. If a worthwhile item has no navigable public link (common for LinkedIn SDUI posts), keep the item but state plainly that no public permalink exists rather than shipping a dead link. Login-walled platforms (x.com / linkedin.com / reddit.com / bsky.app / youtube.com) are exempt from the liveness probe but must still pass the shape check. - Number items sequentially across all sections. 5. Save each topic's report to `reports/{YYYY-MM-DD-HHmm}-{slug}-content.md` (or `reports/{YYYY-MM-DD-HHmm}-content.md` if only one topic). **Exactly one report file per scan.** Every item from every layer (browser-scan sidecars, cascade fallbacks, RSS, APIs, MCP, manual imports) goes into that **one** file β€” never write a separate "supplemental", "addendum", or "sidecar report" alongside it. If a re-scan happens for the same window, edit the existing report file in place. See "One Scan = One Report" in the agent definition for the full rule. diff --git a/CHANGELOG.md b/CHANGELOG.md index 42e9a19..31f9648 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -4,6 +4,35 @@ All notable changes to Content Scout are tracked here. This project uses a product changelog version stream until formal release tags are cut. Minor feature releases use `0.x.0`; major fix bundles also receive their own `0.x.0` entry so every important fix has a durable version number. +## [0.31.0] - 2026-07-02 + +Competitor queries wired into the browser-scan conversations layer. + +### Versioned Features and Fixes + +| Version | Type | Area | Change | +| --- | --- | --- | --- | +| 0.31.0 | Minor feature | Browser scan / Competitor pass | New shared `tools/lib/competitors.mjs` parses the config's `## Competitors` section (bold name + `Aliases:`) into structured entries and provides query-term building + word-boundary matching/tagging. The browser-scan now runs a **competitor pass** over Reddit + X when competitors are configured: it queries the competitor names/aliases, keeps only items that name a tracked competitor, tags each with the matched `competitor`, and writes a `{stamp}-competitors.json` sidecar (plus a `competitors` block in `{stamp}-meta.json`). Because the browser-scan runs as Step 0 of every `/scout-scan`, the competitor pass happens **automatically alongside the monthly mindshare run** β€” no separate command. `/scout-scan` + agent docs ingest the sidecar into the `## Competitor & Market Signals` section and add Hacker News / Stack Overflow / Bluesky per-alias API coverage. Flags: `--no-competitors`, `--max-competitor-terms N` (default 12). | + +### Validation + +- New `tools/web-ui/test/competitors.test.js` (6 tests): parsing, alias handling, query-term building, word-boundary matching (no "Atlas" β†’ "Atlassian" false positive), and item tagging. +- `node --check` on the browser-scan orchestrator + a live `loadConfig` parse of the real config (8 competitors) confirm the wiring. + +## [0.30.0] - 2026-07-02 + +Competitor sentiment & market-signal tracking. + +### Versioned Features and Fixes + +| Version | Type | Area | Change | +| --- | --- | --- | --- | +| 0.30.0 | Minor feature | Reports / Competitor signals | `/scout-scan` now populates a `## Competitor & Market Signals` section when **Competitor tracking** is on and the config's `## Competitors` section lists products: per-competitor content volume, community sentiment (toward the competitor), switching signals (migrations to/from, scored from our product's perspective), and notable announcements β€” matched by competitor name + aliases across conversation sources and the competitor's own blog/release-notes, inline in the single content report. The web UI **Reports** view gains a **Competitors** tab that lists standalone `-competitors.md` deep-dive reports and slices the Competitor & Market Signals section out of content reports (mirroring the Mindshare / CFPs & Events tabs). `tools/lib/doc-meta.mjs` classifies `-competitors.md` as kind "Competitors" and detects the section for tab filtering. | + +### Validation + +- Report classification + section detection verified against a generated `-competitors.md` report and a content report carrying the section. + ## [0.29.0] - 2026-07-02 Monthly roundup, content-originality review, responsive navigation, and scan-source hardening (`feature/roundup-techcommunity-scan`). diff --git a/docs/WORKFLOW.md b/docs/WORKFLOW.md index e443076..f63ee5a 100644 --- a/docs/WORKFLOW.md +++ b/docs/WORKFLOW.md @@ -143,6 +143,12 @@ Turn on **Originality scoring** in your config and each `/scout-scan` adds an `# It is a transparent review aid, **not** a verdict: a low score never drops or down-ranks an item, and titles/snippets return "insufficient text" rather than a guess. For a one-off check outside a scan, use `/scout-originality` (below). +### Competitor Signals (optional) + +Turn on **Competitor tracking** and list rivals under `## Competitors` in your config (each a bold name plus optional `Aliases:`), and every `/scout-scan` adds a `## Competitor & Market Signals` section. The browser-scan Layer 0 runs a **competitor pass** over Reddit and X β€” writing a `{stamp}-competitors.json` sidecar with each mention tagged by the matched competitor β€” and the agent adds Hacker News, Stack Overflow, Bluesky, and each rival's own blog/release-notes. Each competitor gets a row scoring content volume, sentiment *toward that competitor*, switching signals (migrations to/from, from your product's perspective), and notable announcements. + +Because it runs inside the normal scan, a **monthly competitor pass happens automatically alongside your monthly mindshare** β€” there's no separate command. The web UI's **Reports β†’ Competitors** tab surfaces the section (and any standalone deep-dive `-competitors.md` reports). + ### Conversation Tracking Forums and social platforms are scanned separately from blog/article content. Conversations are tracked but not promoted as report items: diff --git a/tools/browser-scan/README.md b/tools/browser-scan/README.md index f80463c..1fe7fae 100644 --- a/tools/browser-scan/README.md +++ b/tools/browser-scan/README.md @@ -220,6 +220,36 @@ Community needs you signed in (see setup); the other three need only a real browser. Hashnode posts on fully custom domains can't be pattern-matched here β€” the open-web Brave layer already covers those. +### Competitor pass (one sidecar, tagged by competitor) + +When the loaded config has `Competitor tracking` set to `on` and a +`## Competitors` section, the scanner runs an extra **competitor pass** over +Reddit and X after the main passes. It queries +the configured competitor names + aliases (built by `competitorQueryTerms` in +`tools/lib/competitors.mjs`), keeps only items that actually name a tracked +competitor, and tags each with the matched competitor. Results merge into one +`*-competitors.json` sidecar, shaped like the platform items above plus these +fields: + +| Field | Meaning | +|---|---| +| `competitor` | The primary tracked competitor the item mentions (canonical name from config) | +| `competitorMatches` | All tracked competitors the item mentions | +| `competitorSentiments` | Target-scoped sentiment and confidence for every matched competitor | +| `competitorSentiment` / `competitorSentimentConfidence` | Convenience fields for the primary match | +| `switchingDirection` | `competitor_to_primary`, `primary_to_competitor`, `competitor_to_competitor`, or `none` | +| `timestamp` / `url` | Normalized scan timestamp and original permalink | + +The `{stamp}-meta.json` sidecar also gains a `competitors` block β€” +`{ names, queryTerms, platforms, mentions, byCompetitor, bySource, +sourceFailures }` β€” for isolated sentiment/mention aggregates and non-fatal +source errors. Competitor sentiment is never included in primary-product +sentiment totals. The agent folds these into the report's **Competitor & Market +Signals** section (never the content sections or Mindshare). Disable with +`--no-competitors`; cap the query set with `--max-competitor-terms N` +(default 12). Hacker News / Stack Overflow / Bluesky competitor coverage runs +in the agent's API layer, not here. + ## Rate-limit hygiene - One in-flight tab per platform; β‰₯3s between page loads. diff --git a/tools/browser-scan/index.mjs b/tools/browser-scan/index.mjs index 795af87..95895c4 100644 --- a/tools/browser-scan/index.mjs +++ b/tools/browser-scan/index.mjs @@ -31,6 +31,7 @@ import { loadConfig } from './lib/config.mjs'; import { ensureProfileDir, launchEdge, attachEdge, newPage } from './lib/browser.mjs'; import { filterHiring, categorizeRoles, ROLE_ORDER } from './lib/hiring-filter.mjs'; import { browserScanSlugDir } from '../lib/paths.mjs'; +import { analyzeCompetitorSources, competitorQueryTerms } from '../lib/competitors.mjs'; const __dirname = path.dirname(fileURLToPath(import.meta.url)); const ROOT = path.resolve(__dirname, '..', '..'); @@ -67,9 +68,13 @@ Usage: [--port 9222] (cdp port) [--days 30] [--max-per-term 25] [--headed] [--since YYYY-MM-DD] [--until YYYY-MM-DD] + [--no-competitors] [--max-competitor-terms 12] --since/--until pin the scan to an exact date range (e.g. one calendar month); the Google Web pass maps it to a precise tbs=cdr:1,cd_min:…,cd_max:… filter. Default is a rolling --days window. + When the config has a "## Competitors" section, a competitor pass runs + over Reddit + X and writes a {stamp}-competitors.json sidecar + (--no-competitors skips it). node index.mjs login --platform x|linkedin|reddit|google|content-sites (LEGACY launch-mode only) @@ -180,7 +185,10 @@ if (command === 'launch') { const windowNote = (flags.since || flags.until) ? `${new Date(sinceMs).toISOString().slice(0, 10)}…${new Date(untilMs).toISOString().slice(0, 10)}` : `${days}d window`; - console.log(`[browser-scan] Loaded config "${slug}" β€” ${config.searchTerms.length} search terms, ${windowNote}, mode=${mode}`); + const competitorNote = Array.isArray(config.competitors) && config.competitors.length + ? `, ${config.competitors.length} competitors` + : ''; + console.log(`[browser-scan] Loaded config "${slug}" β€” ${config.searchTerms.length} search terms${competitorNote}, ${windowNote}, mode=${mode}`); const stamp = formatStamp(new Date()); const outDir = browserScanSlugDir(slug); @@ -289,6 +297,79 @@ if (command === 'launch') { console.log(`[browser-scan] ${platform}: ${kept.length} items${droppedNote} β†’ ${path.relative(ROOT, outFile)}`); } + // ---- competitor pass (conversations layer, filtered by config `## Competitors`) ---- + // Additive + defensive: reuses the same logged-in tab and the same platform + // scanners, but queries the configured competitor names/aliases instead of the + // product terms, tags each hit with the matched competitor, and writes a + // combined {stamp}-competitors.json sidecar that the agent folds into the + // report's Competitor & Market Signals section. Gated on competitors being + // configured; disable with --no-competitors. Wrapped so a failure here never + // affects the product sidecars written above. + let metaCompetitors = null; + const competitorsEnabled = !flags['no-competitors'] && config.competitorTracking + && Array.isArray(config.competitors) && config.competitors.length > 0; + if (competitorsEnabled) { + // Only the conversation platforms browser-scan covers. HN / Stack Overflow / + // Bluesky competitor queries run in the agent's API layer (see scout-scan). + const convoPlatforms = requested.filter((p) => p === 'reddit' || p === 'x'); + const compTerms = competitorQueryTerms(config.competitors, { + max: Number(flags['max-competitor-terms'] || 12), + }); + if (convoPlatforms.length && compTerms.length) { + console.log(`[browser-scan] competitor pass β€” ${config.competitors.length} competitors, ${compTerms.length} query terms over ${convoPlatforms.join(', ')}`); + const competitorSources = []; + for (const platform of convoPlatforms) { + let handle = sharedHandle; + let ownsHandle = false; + try { + if (mode === 'launch') { + const profileDir = ensureProfileDir(__dirname, platform); + if (!hasSession(profileDir)) { + competitorSources.push({ source: platform, items: [] }); + continue; + } + handle = await launchEdge({ profileDir, headed }); + ownsHandle = true; + } + const ctx = { + searchTerms: compTerms, + sinceMs, + untilMs, + maxPerTerm: Math.min(maxPerTerm, 15), + slug, + outDir, + page: mode === 'cdp' ? sharedPage : undefined, + }; + let items = []; + if (platform === 'x') items = await scanX(handle, ctx); + else if (platform === 'reddit') items = await scanReddit(handle, ctx); + for (const it of items) { it.platform = platform; } + competitorSources.push({ source: platform, items: filterHiring(items).kept }); + } catch (e) { + console.error(`[browser-scan] competitor/${platform}: error β€” ${e.message}`); + competitorSources.push({ source: platform, error: e }); + } finally { + if (ownsHandle) await handle.browser.close().catch(() => {}); + } + } + // Keep only items that actually name a tracked competitor; tag each with + // the matched competitor(s). Then drop hiring/recruiting as elsewhere. + const analyzed = analyzeCompetitorSources(competitorSources, config.competitors, { + primaryProduct: [config.primaryProduct, ...config.searchTerms].filter(Boolean), + }); + const compFile = path.join(outDir, `${stamp}-competitors.json`); + fs.writeFileSync(compFile, JSON.stringify(analyzed.items, null, 2)); + metaCompetitors = { + names: config.competitors.map((c) => c.name), + queryTerms: compTerms.length, + platforms: convoPlatforms, + ...analyzed.aggregates, + sourceFailures: analyzed.sourceFailures, + }; + console.log(`[browser-scan] competitors: ${analyzed.items.length} tagged mentions β†’ ${path.relative(ROOT, compFile)}`); + } + } + // Write the meta sidecar (always, even when zero drops, so absence of the // file means "no scan ran" rather than "scan ran but zeros"). const totalDropped = Object.values(hiringDropped).reduce((a, b) => a + b, 0); @@ -302,6 +383,7 @@ if (command === 'launch') { hiringDroppedByMonth, hiringRolesByMonth, hiringRoleOrder: ROLE_ORDER, + competitors: metaCompetitors, }, null, 2)); // In CDP mode we do NOT close the user's Edge β€” they own it. Just diff --git a/tools/browser-scan/lib/config.mjs b/tools/browser-scan/lib/config.mjs index 35d68ef..5ab3eba 100644 --- a/tools/browser-scan/lib/config.mjs +++ b/tools/browser-scan/lib/config.mjs @@ -6,6 +6,7 @@ import fs from 'node:fs'; import path from 'node:path'; import { CONFIGS_DIR } from '../../lib/paths.mjs'; +import { competitorTrackingEnabled, parseCompetitors } from '../../lib/competitors.mjs'; export function loadConfig(root, slug) { // Config moved from .github/prompts/scout-config-{slug}.prompt.md to the @@ -34,9 +35,16 @@ export function loadConfig(root, slug) { path: configPath, raw, searchTerms, + competitorTracking: competitorTrackingEnabled(raw), + competitors: parseCompetitors(raw), + primaryProduct: extractProductName(raw), }; } +function extractProductName(md) { + return (String(md || '').match(/^#\s+scout-config:\s*(.+?)\s*$/mi) || [])[1] || ''; +} + function extractSearchTerms(md) { // Look for a "## Search Terms" section followed by bullet lines. // Falls back to "## Text Search Terms" / "## Hashtags" headers. diff --git a/tools/lib/competitors.mjs b/tools/lib/competitors.mjs new file mode 100644 index 0000000..58ceac3 --- /dev/null +++ b/tools/lib/competitors.mjs @@ -0,0 +1,313 @@ +// competitors.mjs β€” shared parsing + matching for the `## Competitors` +// section of a scout-config file. Pure and dependency-free so both the +// browser-scan conversations layer and the web UI can reuse it. +// +// Config format (one bullet per competitor, bold name + optional Aliases clause): +// +// ## Competitors +// +// - **Amazon DynamoDB** β€” AWS managed NoSQL. Aliases: DynamoDB, DDB, Dynamo. +// - **MongoDB Atlas** β€” managed document DB. Aliases: MongoDB, Atlas, Mongo. +// +// A `_None tracked…_` placeholder (or a missing section) yields an empty list. + +// Pull the raw text of the `## Competitors` section (up to the next `## `). +function extractCompetitorsSection(raw) { + const text = String(raw || ''); + // NOTE: terminate with `(?![\s\S])` for end-of-string β€” JS has no `\Z`, and + // under the /i flag a literal `\Z` degrades to a case-insensitive "z" match + // (which would truncate the section at the "z" in "Azure"). + const m = text.match(/^##\s+Competitors\b[^\n]*\n([\s\S]*?)(?=^##\s|(?![\s\S]))/mi); + return m ? m[1] : ''; +} + +// Strip a trailing "(parenthetical)" from a competitor name, e.g. +// "DataStax Astra DB (Apache Cassandra)" β†’ "DataStax Astra DB". +function stripParenthetical(name) { + return String(name || '').replace(/\s*\([^)]*\)\s*$/, '').trim(); +} + +// Parse the `## Competitors` section into structured entries. +// Returns: [{ name, aliases: [...] }] where `aliases` always includes the +// canonical name (and the name without any trailing parenthetical). Names and +// aliases are de-duplicated case-insensitively while preserving first spelling. +export function parseCompetitors(raw) { + const section = extractCompetitorsSection(raw); + if (!section || /_?None tracked/i.test(section)) return []; + const out = []; + for (const line of section.split(/\r?\n/)) { + const bullet = line.match(/^\s*[-*+]\s+(.+?)\s*$/); + if (!bullet) continue; // skip intro prose + the italic "_Adjacent…_" line + const body = bullet[1]; + // Name: prefer the **bold** span; else the text before an em/en dash, + // colon, or "(". + const boldMatch = body.match(/\*\*(.+?)\*\*/); + let name = boldMatch ? boldMatch[1].trim() : body.split(/[—–:(]/)[0].trim(); + name = name.replace(/^["']|["']$/g, '').trim(); + if (!name) continue; + // Aliases: everything after an "Aliases:" label, comma-separated, up to + // the sentence's terminating period. + const aliasMatch = body.match(/Aliases?\s*:\s*([^.]+?)\s*\.?\s*$/i) + || body.match(/\(\s*aliases?\s*:\s*([^)]+)\)/i); + const aliasList = aliasMatch + ? aliasMatch[1].split(',').map((a) => a.trim()).filter(Boolean) + : []; + const aliasSet = []; + const seen = new Set(); + for (const a of [name, stripParenthetical(name), ...aliasList]) { + const key = a.toLowerCase(); + if (a && !seen.has(key)) { seen.add(key); aliasSet.push(a); } + } + out.push({ name, aliases: aliasSet }); + } + return out; +} + +export function competitorTrackingEnabled(raw) { + const match = String(raw || '').match(/-\s*\*\*Competitor tracking:\*\*\s*(on|off)\b/i); + return match ? match[1].toLowerCase() === 'on' : false; +} + +// Build a flat, de-duplicated list of search-query terms for the conversations +// layer. Prefers the distinctive full names first (less noisy than short +// aliases like "Atlas" or "DDB"), then fills with aliases up to `max`. +export function competitorQueryTerms(competitors, { max = 12, includeAliases = true } = {}) { + const terms = []; + const seen = new Set(); + const push = (t) => { + const v = String(t || '').trim(); + const key = v.toLowerCase(); + if (v && !seen.has(key)) { seen.add(key); terms.push(v); } + }; + // Pass 1: canonical names (most distinctive). + for (const c of competitors || []) push(stripParenthetical(c.name)); + // Pass 2: remaining aliases, if requested and budget remains. + if (includeAliases) { + for (const c of competitors || []) { + for (const a of c.aliases || []) push(a); + } + } + return terms.slice(0, Math.max(0, max)); +} + +// Escape a string for use inside a RegExp. +function escapeRe(s) { + return String(s).replace(/[.*+?^${}()|[\]\\]/g, '\\$&'); +} + +// Return the names of every competitor whose name or any alias appears in +// `text` (case-insensitive, word-boundary matched so "Atlas" doesn't match +// inside "Atlassian"). +export function matchCompetitors(text, competitors) { + const hay = String(text || ''); + if (!hay.trim()) return []; + const hits = []; + for (const c of competitors || []) { + const matched = (c.aliases || []).some((alias) => { + const a = String(alias || '').trim(); + if (!a) return false; + // \b around the escaped alias; the alias may contain spaces (phrase). + const re = new RegExp(`(^|[^\\w])${escapeRe(a)}(?=$|[^\\w])`, 'i'); + return re.test(hay); + }); + if (matched) hits.push(c.name); + } + return hits; +} + +// Collect the text fields a browser-scan item might carry, for matching. +function itemHaystack(item) { + if (!item || typeof item !== 'object') return ''; + const fields = ['title', 'text', 'selftext', 'body', 'snippet', 'excerpt', 'description']; + return fields.map((f) => item[f]).filter((v) => typeof v === 'string').join(' \u2014 '); +} + +// Tag each item with the competitor(s) it mentions and drop items that mention +// none. Returns a new array; original items are shallow-cloned with two added +// fields: `competitor` (first match) and `competitorMatches` (all matches). +export function tagCompetitorItems(items, competitors) { + const out = []; + for (const item of items || []) { + const matches = matchCompetitors(itemHaystack(item), competitors); + if (!matches.length) continue; + out.push({ ...item, competitor: matches[0], competitorMatches: matches }); + } + return out; +} + +function itemUrl(item) { + return String(item?.url || item?.permalink || item?.link || '').trim(); +} + +export function canonicalCompetitorUrl(url) { + try { + const parsed = new URL(String(url || '').trim()); + parsed.hash = ''; + parsed.hostname = parsed.hostname.toLowerCase() + .replace(/^www\./, '') + .replace(/^old\.reddit\.com$/, 'reddit.com') + .replace(/^twitter\.com$/, 'x.com'); + for (const key of [...parsed.searchParams.keys()]) { + if (/^(utm_.+|ref|source|fbclid|gclid)$/i.test(key)) parsed.searchParams.delete(key); + } + parsed.pathname = parsed.pathname.replace(/\/+$/, '') || '/'; + return parsed.toString(); + } catch { + return String(url || '').trim().replace(/[?#].*$/, '').replace(/\/+$/, '').toLowerCase(); + } +} + +export function dedupeCompetitorItems(items) { + const output = []; + const byUrl = new Map(); + for (const item of items || []) { + const key = canonicalCompetitorUrl(itemUrl(item)); + if (!key) { + output.push({ ...item }); + continue; + } + const existing = byUrl.get(key); + if (!existing) { + const copy = { ...item, canonicalUrl: key }; + byUrl.set(key, copy); + output.push(copy); + continue; + } + existing.competitorMatches = [...new Set([ + ...(existing.competitorMatches || []), + ...(item.competitorMatches || []), + ])]; + } + return output; +} + +function aliasesFor(name, competitors) { + const competitor = (competitors || []).find((entry) => entry.name === name); + return competitor?.aliases?.length ? competitor.aliases : [name]; +} + +function mentionsAny(text, terms) { + return (terms || []).some((term) => { + const value = String(term || '').trim(); + return value && new RegExp(`(^|[^\\w])${escapeRe(value)}(?=$|[^\\w])`, 'i').test(text); + }); +} + +export function detectSwitchingDirection(text, primaryProduct, competitors) { + const hay = String(text || ''); + const migration = hay.match(/\b(?:migrat(?:e|ed|ing)|mov(?:e|ed|ing)|switch(?:ed|ing)?|transition(?:ed|ing)?)\s+from\s+(.{1,100}?)\s+to\s+(.{1,100}?)(?:[.!?]|$)/i); + if (!migration) return 'none'; + const [, from, to] = migration; + const primaryTerms = Array.isArray(primaryProduct) ? primaryProduct : [primaryProduct]; + const fromPrimary = mentionsAny(from, primaryTerms); + const toPrimary = mentionsAny(to, primaryTerms); + const fromCompetitor = matchCompetitors(from, competitors).length > 0; + const toCompetitor = matchCompetitors(to, competitors).length > 0; + if (fromCompetitor && toPrimary) return 'competitor_to_primary'; + if (fromPrimary && toCompetitor) return 'primary_to_competitor'; + if (fromCompetitor && toCompetitor) return 'competitor_to_competitor'; + return 'none'; +} + +function scopedSentences(text, aliases) { + return String(text || '') + .split(/(?<=[.!?])\s+|\n+/) + .filter((sentence) => mentionsAny(sentence, aliases)) + .join(' '); +} + +export function classifyCompetitorSentiment(text, competitorName, competitors, switchingDirection = 'none') { + const scoped = scopedSentences(text, aliasesFor(competitorName, competitors)); + const positive = /\b(?:love|loved|great|excellent|reliable|recommend(?:ed)?|fast|better|best|impressed|happy|satisfied)\b/i.test(scoped); + const negative = /\b(?:hate|hated|bad|awful|unreliable|slow|worse|worst|expensive|frustrat(?:ed|ing)|outage|buggy|abandon(?:ed|ing)|left)\b/i.test(scoped); + if (positive && negative) return { sentiment: 'mixed', confidence: 'high' }; + if (positive) return { sentiment: 'positive', confidence: 'high' }; + if (negative) return { sentiment: 'negative', confidence: 'high' }; + if (switchingDirection === 'competitor_to_competitor') { + return { sentiment: 'neutral', confidence: 'low' }; + } + return { sentiment: 'neutral', confidence: scoped ? 'medium' : 'low' }; +} + +function emptySentiments() { + return { positive: 0, neutral: 0, negative: 0, mixed: 0, unknown: 0 }; +} + +export function analyzeCompetitorSources(sourceResults, competitors, { + primaryProduct = '', + classify = classifyCompetitorSentiment, +} = {}) { + const sourceFailures = []; + const collected = []; + for (const result of sourceResults || []) { + const source = String(result?.source || result?.platform || 'unknown'); + if (result?.error) { + sourceFailures.push({ source, error: String(result.error.message || result.error) }); + continue; + } + for (const item of result?.items || []) collected.push({ ...item, platform: item.platform || source }); + } + + const tagged = tagCompetitorItems(collected, competitors); + const deduped = dedupeCompetitorItems(tagged); + const items = deduped.map((item) => { + const text = itemHaystack(item); + const switchingDirection = detectSwitchingDirection(text, primaryProduct, competitors); + const competitorSentiments = (item.competitorMatches || []).map((name) => { + try { + const verdict = classify(text, name, competitors, switchingDirection) || {}; + return { + competitor: name, + sentiment: ['positive', 'neutral', 'negative', 'mixed', 'unknown'].includes(verdict.sentiment) + ? verdict.sentiment + : 'unknown', + confidence: ['high', 'medium', 'low'].includes(verdict.confidence) + ? verdict.confidence + : 'low', + }; + } catch { + return { competitor: name, sentiment: 'unknown', confidence: 'low' }; + } + }); + const first = competitorSentiments[0] || { + competitor: item.competitor, + sentiment: 'unknown', + confidence: 'low', + }; + return { + ...item, + url: itemUrl(item), + timestamp: item.timestamp || item.post_date || item.date || item.published_at || '', + competitorSentiments, + competitorSentiment: first.sentiment, + competitorSentimentConfidence: first.confidence, + switchingDirection, + }; + }); + + const byCompetitor = {}; + const bySource = {}; + for (const item of items) { + const source = item.platform || 'unknown'; + bySource[source] ||= { mentions: 0, sentiments: emptySentiments() }; + bySource[source].mentions += 1; + for (const verdict of item.competitorSentiments) { + byCompetitor[verdict.competitor] ||= { + mentions: 0, + sentiments: emptySentiments(), + bySource: {}, + }; + const aggregate = byCompetitor[verdict.competitor]; + aggregate.mentions += 1; + aggregate.sentiments[verdict.sentiment] += 1; + aggregate.bySource[source] = (aggregate.bySource[source] || 0) + 1; + bySource[source].sentiments[verdict.sentiment] += 1; + } + } + + return { + items, + aggregates: { mentions: items.length, byCompetitor, bySource }, + sourceFailures, + }; +} diff --git a/tools/lib/doc-meta.mjs b/tools/lib/doc-meta.mjs index 120bfb9..1c09d24 100644 --- a/tools/lib/doc-meta.mjs +++ b/tools/lib/doc-meta.mjs @@ -25,6 +25,7 @@ const KIND_TABLE = [ { match: /-originality\.md$/i, id: 'originality', label: 'Originality' }, { match: /-cfps?\.md$/i, id: 'cfp', label: 'CFPs' }, { match: /-conferences?\.md$/i, id: 'conference', label: 'Conference' }, + { match: /-competitors?\.md$/i, id: 'competitors', label: 'Competitors' }, ]; export function detectKind(name) { @@ -35,7 +36,7 @@ export function detectKind(name) { // Filenames look like: 2026-05-21-1004-azure-cosmos-db-content.md // Capture the stamp + slug so the UI can group runs and surface the // subject independent of the wordy H1 title. -const FILENAME_RE = /^(\d{4}-\d{2}-\d{2})-(\d{4})-(.+?)-(content|mindshare|supplemental|cfps?|conferences?|social-posts|posting-calendar|alt-.+|solo[-.].+|seo[-.].+)\.md$/i; +const FILENAME_RE = /^(\d{4}-\d{2}-\d{2})-(\d{4})-(.+?)-(content|mindshare|supplemental|competitors?|cfps?|conferences?|social-posts|posting-calendar|alt-.+|solo[-.].+|seo[-.].+)\.md$/i; export function parseFilename(name) { const m = String(name || '').match(FILENAME_RE); @@ -129,6 +130,7 @@ function extractDateRange(raw) { const SECTION_MATCHERS = { mindshare: /^#{2,3}\s+mindshare\b/im, cfp: /^#{2,3}\s+(open calls for papers|cfps?\b|calls? for papers\b|conferences?\b|conference content\b)/im, + competitors: /^#{2,3}\s+(competitor|competitive)\b/im, }; function extractSectionFlags(raw) { @@ -136,6 +138,7 @@ function extractSectionFlags(raw) { return { mindshare: SECTION_MATCHERS.mindshare.test(text), cfp: SECTION_MATCHERS.cfp.test(text), + competitors: SECTION_MATCHERS.competitors.test(text), }; } diff --git a/tools/lib/report-index.mjs b/tools/lib/report-index.mjs index 51fdc55..a7f53f4 100644 --- a/tools/lib/report-index.mjs +++ b/tools/lib/report-index.mjs @@ -22,7 +22,8 @@ // // The return shape of parseReport / parseReportFromJson is identical so // callers can treat them interchangeably: -// { slug, generatedAt, items, conversations, sentimentTotals, skippedSources } +// { slug, generatedAt, items, conversations, sentimentTotals, skippedSources, +// competitorMentions, competitorAggregates, competitorSourceFailures } import fs from 'node:fs/promises'; import path from 'node:path'; @@ -39,6 +40,27 @@ function emptySentimentTotals() { return { positive: 0, neutral: 0, negative: 0, mixed: 0, unknown: 0 }; } +function emptyCompetitorAggregates() { + return { mentions: 0, byCompetitor: {}, bySource: {} }; +} + +function normalizeCompetitorAggregates(value) { + if (!value || typeof value !== 'object' || Array.isArray(value)) { + return emptyCompetitorAggregates(); + } + return { + mentions: Number.isFinite(value.mentions) ? value.mentions : 0, + byCompetitor: + value.byCompetitor && typeof value.byCompetitor === 'object' && !Array.isArray(value.byCompetitor) + ? value.byCompetitor + : {}, + bySource: + value.bySource && typeof value.bySource === 'object' && !Array.isArray(value.bySource) + ? value.bySource + : {}, + }; +} + function normalizeName(name) { return String(name || '') .toLowerCase() @@ -634,6 +656,9 @@ export function parseReportFromJson(rawOrObj, fileName, options = {}) { conversations: filteredConversations, sentimentTotals: sentimentTotalsFor(filteredConversations), skippedSources, + competitorMentions: Array.isArray(data.competitor_mentions) ? data.competitor_mentions : [], + competitorAggregates: normalizeCompetitorAggregates(data.competitor_aggregates), + competitorSourceFailures: Array.isArray(data.competitor_source_failures) ? data.competitor_source_failures : [], }; } @@ -833,6 +858,9 @@ export function parseReport(raw, fileName, options = {}) { conversations: filteredConversations, sentimentTotals: sentimentTotalsFor(filteredConversations), skippedSources, + competitorMentions: [], + competitorAggregates: emptyCompetitorAggregates(), + competitorSourceFailures: [], }; } diff --git a/tools/web-ui/public/index.html b/tools/web-ui/public/index.html index a978d66..6557801 100644 --- a/tools/web-ui/public/index.html +++ b/tools/web-ui/public/index.html @@ -1018,6 +1018,7 @@

Reports

+ diff --git a/tools/web-ui/public/pages/reports.js b/tools/web-ui/public/pages/reports.js index 5aa3796..de612f6 100644 --- a/tools/web-ui/public/pages/reports.js +++ b/tools/web-ui/public/pages/reports.js @@ -2,10 +2,11 @@ import { $, api, escape, escapeAttr } from '../lib/core.js'; import { wireListFilter, renderDocListItem, renderDocBody } from '../lib/doc-list.js'; import { setReportsPayload } from './report-state.js'; -const REPORTS_TABS = ['content', 'mindshare', 'cfp', 'roundup']; +const REPORTS_TABS = ['content', 'mindshare', 'cfp', 'roundup', 'competitors']; const TAB_SECTION_MATCH = { mindshare: /^mindshare\b/i, cfp: /^(open calls for papers|cfps?\b|calls? for papers\b|conferences?\b)/i, + competitors: /^(competitor|competitive)\b/i, }; const REPORT_SECTIONS = [ { key: 'all', label: 'All', match: null }, @@ -28,6 +29,9 @@ function rowMatchesTab(li, tab) { || (kind === 'content' && li.dataset.hasCfp === '1'); case 'roundup': return kind === 'roundup'; + case 'competitors': + return kind === 'competitors' + || (kind === 'content' && li.dataset.hasCompetitors === '1'); default: return false; } @@ -56,6 +60,9 @@ function setReportsRowTabBadge(li, tab, match) { } else if (tab === 'roundup' && li.dataset.kind === 'roundup') { label = 'Roundup'; kindClass = 'kind-roundup'; + } else if (tab === 'competitors' && ['content', 'competitors'].includes(li.dataset.kind || '')) { + label = 'Competitors'; + kindClass = 'kind-competitors'; } badge.textContent = label; for (const cls of [...badge.classList]) { @@ -157,14 +164,16 @@ async function openReportRow(li) { if (sectionRe && kind === 'content') { const sliced = extractSectionsHtml(report.html, sectionRe); if (sliced) { - const label = reportsActiveTab === 'cfp' ? 'CFPs & Events' : 'Mindshare'; + const label = reportsActiveTab === 'cfp' ? 'CFPs & Events' + : reportsActiveTab === 'competitors' ? 'Competitor & Market Signals' + : 'Mindshare'; renderDocBody(body, { name, html: `

${escape(label)} section of ${escape(name)} β€” open the Full Report tab for the full report.

${sliced}`, kind: 'reports', }); } else { - body.innerHTML = `

This scan report has no ${reportsActiveTab === 'cfp' ? 'CFP/Conferences' : 'Mindshare'} section. Open the Full Report tab for the full report.

`; + body.innerHTML = `

This scan report has no ${reportsActiveTab === 'cfp' ? 'CFP/Conferences' : reportsActiveTab === 'competitors' ? 'Competitor' : 'Mindshare'} section. Open the Full Report tab for the full report.

`; } } else { renderDocBody(body, { name, html: report.html, kind: 'reports' }); @@ -239,6 +248,8 @@ function renderReportsActionBar(tab) { bar.innerHTML = `CFPs & Events only β€” the Calls for Papers and Conferences sections of each scan report.`; } else if (tab === 'roundup') { renderRoundupActionBar(bar); + } else if (tab === 'competitors') { + bar.innerHTML = `Competitor signals only β€” the Competitor & Market Signals section of each scan report, plus any standalone competitor reports. Tracks the products listed under ## Competitors in your config.`; } else { bar.innerHTML = ''; } @@ -350,6 +361,7 @@ export async function loadReports() { const sections = meta.sections || {}; li.dataset.hasMindshare = sections.mindshare ? '1' : ''; li.dataset.hasCfp = sections.cfp ? '1' : ''; + li.dataset.hasCompetitors = sections.competitors ? '1' : ''; li.addEventListener('click', (event) => { if (event.target.closest('.entry-open')) return; openReportRow(li); diff --git a/tools/web-ui/public/theme-modern.css b/tools/web-ui/public/theme-modern.css index 69d5ea3..b2a9549 100644 --- a/tools/web-ui/public/theme-modern.css +++ b/tools/web-ui/public/theme-modern.css @@ -3444,4 +3444,5 @@ a.help-dot { /* Kind-badge colors for new doc kinds */ aside li[data-name] .entry-kind.kind-cfp { background: rgba(14,165,233,0.16); color: #bae6fd; border-color: rgba(14,165,233,0.35); } -aside li[data-name] .entry-kind.kind-conference { background: rgba(217,70,239,0.16); color: #f5d0fe; border-color: rgba(217,70,239,0.35); } \ No newline at end of file +aside li[data-name] .entry-kind.kind-conference { background: rgba(217,70,239,0.16); color: #f5d0fe; border-color: rgba(217,70,239,0.35); } +aside li[data-name] .entry-kind.kind-competitors { background: rgba(245,158,11,0.16); color: #fde68a; border-color: rgba(245,158,11,0.35); } \ No newline at end of file diff --git a/tools/web-ui/test/competitors.test.js b/tools/web-ui/test/competitors.test.js new file mode 100644 index 0000000..48a950d --- /dev/null +++ b/tools/web-ui/test/competitors.test.js @@ -0,0 +1,176 @@ +import { test } from 'node:test'; +import assert from 'node:assert/strict'; +import { + analyzeCompetitorSources, + competitorTrackingEnabled, + parseCompetitors, + competitorQueryTerms, + detectSwitchingDirection, + matchCompetitors, + tagCompetitorItems, +} from '../../lib/competitors.mjs'; + +const SAMPLE_CONFIG = `# scout-config: Azure Cosmos DB + +## Competitors + +Closest Azure Cosmos DB competitors β€” tracked for content volume, sentiment, switching signals, and product announcements. Aliases help match mentions across sources. + +- **Amazon DynamoDB** β€” AWS fully-managed serverless NoSQL (key-value + document); the most directly compared alternative. Aliases: DynamoDB, DDB, Dynamo. +- **MongoDB Atlas** β€” managed document database; overlaps Cosmos DB's MongoDB API (RU + vCore). Aliases: MongoDB, Atlas, Mongo, MongoDB Atlas. +- **DataStax Astra DB (Apache Cassandra)** β€” managed Cassandra wide-column NoSQL; overlaps Cosmos DB's Cassandra API. Aliases: Cassandra, Apache Cassandra, Astra, Astra DB, DataStax. +- **ScyllaDB** β€” high-performance Cassandra-compatible NoSQL. Aliases: ScyllaDB, Scylla. + +_Adjacent (watch, not primary): Redis / Azure Managed Redis, PlanetScale, Fauna, Aerospike, TiDB._ + +## Conferences & Events +`; + +test('parseCompetitors extracts bold names + aliases and skips prose/adjacent lines', () => { + const competitors = parseCompetitors(SAMPLE_CONFIG); + assert.equal(competitors.length, 4); + const dynamo = competitors[0]; + assert.equal(dynamo.name, 'Amazon DynamoDB'); + // Name is always part of its own alias set, plus the parsed aliases. + assert.ok(dynamo.aliases.includes('Amazon DynamoDB')); + assert.ok(dynamo.aliases.includes('DynamoDB')); + assert.ok(dynamo.aliases.includes('DDB')); + // The "_Adjacent…_" italic line is not a competitor. + assert.ok(!competitors.some((c) => /Redis|PlanetScale|Aerospike/.test(c.name))); +}); + +test('parseCompetitors keeps the parenthetical name but also matches the stripped form', () => { + const competitors = parseCompetitors(SAMPLE_CONFIG); + const astra = competitors.find((c) => /Astra/.test(c.name)); + assert.equal(astra.name, 'DataStax Astra DB (Apache Cassandra)'); + assert.ok(astra.aliases.includes('DataStax Astra DB')); // parenthetical stripped + assert.ok(astra.aliases.includes('Cassandra')); +}); + +test('parseCompetitors returns [] for "None tracked" and a missing section', () => { + assert.deepEqual(parseCompetitors('## Competitors\n_None tracked. Add to enable._\n'), []); + assert.deepEqual(parseCompetitors('# config\n\n## Topics\n- foo\n'), []); +}); + +test('competitor tracking requires an explicit on toggle', () => { + assert.equal(competitorTrackingEnabled('- **Competitor tracking:** on'), true); + assert.equal(competitorTrackingEnabled('- **Competitor tracking:** off'), false); + assert.equal(competitorTrackingEnabled(SAMPLE_CONFIG), false); +}); + +test('competitorQueryTerms prefers distinctive names first and caps the list', () => { + const competitors = parseCompetitors(SAMPLE_CONFIG); + const terms = competitorQueryTerms(competitors, { max: 4 }); + assert.deepEqual(terms, ['Amazon DynamoDB', 'MongoDB Atlas', 'DataStax Astra DB', 'ScyllaDB']); + // De-dupes case-insensitively across names + aliases. + const all = competitorQueryTerms(competitors, { max: 100 }); + const lower = all.map((t) => t.toLowerCase()); + assert.equal(new Set(lower).size, lower.length); +}); + +test('matchCompetitors word-boundary matches and avoids substring false positives', () => { + const competitors = parseCompetitors(SAMPLE_CONFIG); + assert.deepEqual(matchCompetitors('We migrated from DynamoDB to Cosmos DB', competitors), ['Amazon DynamoDB']); + // "Atlas" alias should NOT fire inside "Atlassian". + assert.deepEqual(matchCompetitors('We use Atlassian Jira', competitors), []); + // Multi-word alias + multiple competitors in one string. + const hits = matchCompetitors('Comparing MongoDB Atlas vs ScyllaDB for scale', competitors); + assert.ok(hits.includes('MongoDB Atlas')); + assert.ok(hits.includes('ScyllaDB')); +}); + +test('tagCompetitorItems tags matches and drops items mentioning no competitor', () => { + const competitors = parseCompetitors(SAMPLE_CONFIG); + const items = [ + { title: 'Why we left DynamoDB', url: 'https://example.com/a' }, + { title: 'A post about Postgres tuning', url: 'https://example.com/b' }, + { title: 'ScyllaDB 2026.2 released', text: 'adds DynamoDB Streams', url: 'https://example.com/c' }, + ]; + const tagged = tagCompetitorItems(items, competitors); + assert.equal(tagged.length, 2); + assert.equal(tagged[0].competitor, 'Amazon DynamoDB'); + // Item c mentions both ScyllaDB (title) and DynamoDB (text). + const c = tagged.find((t) => /2026\.2/.test(t.title)); + assert.ok(c.competitorMatches.includes('ScyllaDB')); + assert.ok(c.competitorMatches.includes('Amazon DynamoDB')); +}); + +test('analyzeCompetitorSources deduplicates canonical URLs before classification', () => { + const competitors = parseCompetitors(SAMPLE_CONFIG); + let classifications = 0; + const result = analyzeCompetitorSources([ + { + source: 'x', + items: [{ text: 'DynamoDB is great', url: 'https://twitter.com/user/status/1?utm_source=test' }], + }, + { + source: 'reddit', + items: [{ text: 'DDB is great', url: 'https://x.com/user/status/1' }], + }, + ], competitors, { + classify: () => { + classifications += 1; + return { sentiment: 'positive', confidence: 'high' }; + }, + }); + + assert.equal(result.items.length, 1); + assert.equal(classifications, 1); + assert.equal(result.aggregates.byCompetitor['Amazon DynamoDB'].mentions, 1); +}); + +test('competitor sentiment is isolated from primary-product sentiment', () => { + const competitors = parseCompetitors(SAMPLE_CONFIG); + const result = analyzeCompetitorSources([{ + source: 'reddit', + items: [{ + title: 'Cosmos DB is awful. DynamoDB is great.', + url: 'https://reddit.com/r/databases/comments/1/example', + }], + }], competitors, { primaryProduct: 'Cosmos DB' }); + + const item = result.items[0]; + assert.equal(item.competitorSentiment, 'positive'); + assert.equal(item.competitorSentimentConfidence, 'high'); + assert.equal(item.switchingDirection, 'none'); +}); + +test('switching direction stays separate from competitor sentiment', () => { + const competitors = parseCompetitors(SAMPLE_CONFIG); + assert.equal( + detectSwitchingDirection('We migrated from DynamoDB to Cosmos DB.', 'Cosmos DB', competitors), + 'competitor_to_primary', + ); + assert.equal( + detectSwitchingDirection('We switched from Cosmos DB to MongoDB Atlas.', 'Cosmos DB', competitors), + 'primary_to_competitor', + ); + + const result = analyzeCompetitorSources([{ + source: 'x', + items: [{ + text: 'We moved from DynamoDB to MongoDB Atlas.', + url: 'https://x.com/user/status/2', + post_date: '2026-07-20T12:00:00Z', + }], + }], competitors, { primaryProduct: 'Cosmos DB' }); + assert.equal(result.items[0].switchingDirection, 'competitor_to_competitor'); + assert.equal(result.items[0].competitorSentiment, 'neutral'); + assert.equal(result.items[0].timestamp, '2026-07-20T12:00:00Z'); +}); + +test('partial competitor-source failure preserves successful results and aggregates', () => { + const competitors = parseCompetitors(SAMPLE_CONFIG); + const result = analyzeCompetitorSources([ + { source: 'reddit', error: new Error('rate limited') }, + { + source: 'x', + items: [{ text: 'ScyllaDB is fast', url: 'https://x.com/user/status/3' }], + }, + ], competitors); + + assert.equal(result.items.length, 1); + assert.deepEqual(result.sourceFailures, [{ source: 'reddit', error: 'rate limited' }]); + assert.equal(result.aggregates.bySource.x.mentions, 1); + assert.equal(result.aggregates.byCompetitor.ScyllaDB.sentiments.positive, 1); +}); diff --git a/tools/web-ui/test/report-index.test.js b/tools/web-ui/test/report-index.test.js index 0b26772..53862b1 100644 --- a/tools/web-ui/test/report-index.test.js +++ b/tools/web-ui/test/report-index.test.js @@ -134,6 +134,87 @@ test('parseReportFromJson drops conversations that duplicate canonical YouTube i assert.equal(parsed.sentimentTotals.positive, 0); }); +test('parseReportFromJson exposes competitor aggregates without changing primary sentiment totals', () => { + const competitorMention = { + competitor: 'Amazon DynamoDB', + competitorSentiment: 'negative', + competitorSentimentConfidence: 'high', + platform: 'reddit', + timestamp: '2026-07-20T12:00:00Z', + url: 'https://reddit.com/r/databases/comments/1/example', + switchingDirection: 'competitor_to_primary', + }; + const competitorAggregates = { + mentions: 1, + byCompetitor: { + 'Amazon DynamoDB': { + mentions: 1, + sentiments: { positive: 0, neutral: 0, negative: 1, mixed: 0, unknown: 0 }, + bySource: { reddit: 1 }, + }, + }, + bySource: { reddit: { mentions: 1 } }, + }; + const parsed = parseReportFromJson({ + generated_at: '2026-07-20', + items: [], + competitor_mentions: [competitorMention], + competitor_aggregates: competitorAggregates, + competitor_source_failures: [{ source: 'x', error: 'unavailable' }], + }, '2026-07-20-1200-test-product-content.md'); + + assert.deepEqual(parsed.competitorMentions, [competitorMention]); + assert.deepEqual(parsed.competitorAggregates, competitorAggregates); + assert.deepEqual(parsed.competitorSourceFailures, [{ source: 'x', error: 'unavailable' }]); + assert.deepEqual(parsed.sentimentTotals, { + positive: 0, + neutral: 0, + negative: 0, + mixed: 0, + unknown: 0, + }); +}); + +test('parseReportFromJson normalizes malformed competitor metadata to stable defaults', () => { + const parsed = parseReportFromJson({ + generated_at: '2026-07-20', + items: [], + competitor_mentions: {}, + competitor_aggregates: 'broken', + competitor_source_failures: { source: 'x', error: 'unavailable' }, + }, '2026-07-20-1200-test-product-content.md'); + + assert.deepEqual(parsed.competitorMentions, []); + assert.deepEqual(parsed.competitorAggregates, { + mentions: 0, + byCompetitor: {}, + bySource: {}, + }); + assert.deepEqual(parsed.competitorSourceFailures, []); +}); + +test('parseReport exposes default competitor fields for legacy markdown reports', () => { + const report = ` +**Generated:** 2026-07-20 + +## Official Content + +| # | Date | Title | Channel | Tags | EP | Link | +|---|------|-------|---------|------|----|------| +| 1 | 2026-07-20 | Example post | Blog | \`#tag\` | 5 | [link](https://example.com/post) | +`; + + const parsed = parseReport(report, '2026-07-20-1200-test-product-content.md'); + + assert.deepEqual(parsed.competitorMentions, []); + assert.deepEqual(parsed.competitorAggregates, { + mentions: 0, + byCompetitor: {}, + bySource: {}, + }); + assert.deepEqual(parsed.competitorSourceFailures, []); +}); + test('parseReport marks official account conversation rows as product-side', () => { const report = ` **Generated:** 2026-05-08