Skip to content

Gateway details page: precomputed overviews, faster windows, no connection leak - #137

Open
bansalayush247 wants to merge 1 commit into
fedimint:masterfrom
bansalayush247:perf/gateway-details-page
Open

bansalayush247 wants to merge 1 commit into
fedimint:masterfrom
bansalayush247:perf/gateway-details-page

Conversation

@bansalayush247

Copy link
Copy Markdown
Member

Summary

Makes the gateway details page fast and current, and stops the gateway fetch from leaking connections. The backend precomputes each federation's gateway list and uptime trend in the background and serves both from one endpoint. Scope is limited to gateway fetching and the gateway details page; it replaces #135, which also carried app-wide frontend changes.

Changes

Backend

  • Connection leak: /config/:invite/gateways built a new ConnectorRegistry twice per request (config download and gateway fetch), and dropped registries keep their guardian connections open. It now reuses the observer's shared registry, like the background tasks already do.
  • New GET /federations/:id/gateways/overview?window= returns the gateway list and the uptime trend computed at the same moment, plus computed_at. ?include=gateways or ?include=trend returns only that part.
  • A background task rebuilds the overviews for every federation, one query at a time, and keeps them in memory: 24h and 7D every 5 minutes, 30D every 30 minutes, 90D every hour. Requests answer in ~1 ms instead of up to 4 s for 90D. Right after a restart, a missing entry is computed on the spot (and not stored, so unknown IDs can't grow the cache).
  • Removed GET /federations/:id/gateways and GET /federations/:id/gateways/uptime-trend (unstable API, only used by our frontend) and the 1h window, which the page never offered and is shorter than the 5-minute poll interval.
  • Activity CTEs are NOT MATERIALIZED: the planner misestimated them as one row and nested-looped ~10k × 10k rows (90D: 55 s → 4 s, same results).
  • A gateway missing from the registry for 7 days counts as retired: later samples no longer count against uptime or the trend. Before, every gateway ever seen stayed at 0% forever.
  • Gateway polls are stored again (a timestamptz parameter was inferred as interval, so every insert failed).

Gateway details page

  • One request per load, refreshed every minute while the tab is visible. "Updated" shows when the server computed the data.
  • Window, status filter and sort live in the URL; the last window is remembered.
  • The live registry lookup runs in the background and never holds up the page.
  • Clear error and retry states, with an error boundary around the chart.
  • Retired gateways are listed separately behind a "Retired" filter.
  • Gateways count as online for 15 minutes after being seen (three polls), so they no longer flip to degraded between polls.
  • Cards on phones; only the needed parts of ECharts are bundled.

Testing

  • cargo clippy, tsc, eslint and npm run build pass.
  • The overview endpoint, the per-window refresh and the page have run on our staging instance since 10 October against live federation data: polls are stored every 5 minutes, /overview answers in ~1–2 ms for every window, and the page loads each window with one request and no console errors.
  • Connection leak, measured locally (server against an empty database, open sockets after a warm-up call and then 10 more calls to /config/:invite/gateways): before 31 → 291 (+260, about 26 per call); with this PR 16 → 17.

Not in this PR

fetch_config_inner (/config/:invite) and try_fetch_meta_inner (/config/:invite/meta) still build a registry per request and leak the same way. They aren't gateway code, so they're left for a separate change.

🤖 Generated with Claude Code

…iews

The gateway page made two requests per load and the 90D window took up to
55 s. A background task now precomputes each federation's gateway list and
uptime trend per window and serves both from
/federations/:id/gateways/overview, so the page loads with one request in
about a millisecond.

- 24h/7D refresh every 5 minutes, 30D every 30, 90D hourly
- activity CTEs are NOT MATERIALIZED (90D: 55 s to 4 s, same results)
- gateways missing from the registry for 7 days count as retired and no
  longer drag uptime down; the page lists them separately
- gateway polls are stored again (the window start parameter was
  inferred as an interval)
- remove the old /gateways and /gateways/uptime-trend endpoints and the
  1h window
- /config/:invite/gateways reuses the observer's connector registry
  instead of building two per request, which left their guardian
  connections open
- the page keeps its view in the URL, refreshes every minute, looks up
  the live registry in the background and counts gateways online for
  15 minutes after they were seen

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant