feat: report provisioning reliability across the fleet - #153
Conversation
ttlogan
left a comment
There was a problem hiding this comment.
Reviewed the fleet reliability report. Each denominator keeps its numerator explicit (first-attempt success, retry recovery, cancellation, unsupported, readiness), and the docs match the code. Retry series identity surviving restarts and bundle changes is the right call, and old history aging into a conservatively unknown origin beats guessing. Ran the new suite: 19/19 pass, including the real setup-failure-to-launch-success path. CI is green. Looks good.
ttlogan
left a comment
There was a problem hiding this comment.
Follow-up on the cuda_probe_compile fix. The two probe failures that split by different log tails now land under one v1:cuda_probe_compile signature, and the regression test reproduces the split then confirms grouping. Ran it: passes. Docs baseline now records the real snapshot (10 attempts, 6/4, 9 llama.cpp / 1 vLLM) with the canary blocked on DNS, which is the honest state. Nothing to add.
ttlogan
left a comment
There was a problem hiding this comment.
Read the matched L4 pilot record and re-ran the numbers. 120% of the 74.554863 s baseline is 89.465836 s, and the 72.521237 s canary lands at 97.27%, under the threshold. First-attempt and cancellation/unsupported denominators stay explicit, and the one-sample caveat is called out instead of over-claimed. Keeping the placement-inclusive recovery cycle (626/636 s) separate from the setup-to-readiness definition is right.
Looks good. No new findings on this docs commit.
ttlogan
left a comment
There was a problem hiding this comment.
Re-read the accepted-targets doc change. The rounding is right: 120% of the 74.554863 s baseline is 89.465836 s, which rounds to the 89.5 s threshold, and the 72.521237 s canary still lands at 97.27%, under it. Status text now just records the operator acceptance instead of pending. No code touched, docs match the measured record. Looks good.
sadsfae
left a comment
There was a problem hiding this comment.
Review summary (multi-team pass)
The feature is well designed: stdlib-only report builder, no node I/O, explicit numerators/denominators, honest evidence gaps, and a real end-to-end recovered-series test plus a regression test for the CUDA probe grouping. Auth (require_admin_auth), route ordering, input validation, and the drill-down links all check out, and docs match the metric definitions. Three should-fix items remain.
Verification (head aa794ff): new reliability suites 20 passed (Python + shipped-JS harness); tests/provisioning 717 passed, 1 skipped, 1 env-only failure (test pins uv 0.12.17, host has 0.12.9); tests/api + tests/frontend 845 passed. Coverage per PR docs 93.72% branch; not re-measured here.
Must-fix
- Global
evicted_attemptscounter decidesseries_origin_knownfor every host, so any single eviction makes all later first-ever attempts unmeasured and the headline first_attempt_success metric silently degrades (log_store.py:168). Per-host eviction bookkeeping; the current test pins the global behavior. datetime-localvalues are browser-local but get:00Z, so non-UTC browsers filter and export a different window than labeled (reliability.js:78). Convert to ISO/UTC before sending.- Success is committed in two writes:
_update_state(COMPLETE)sets stage/ready_at but status stays running until_mark_log_completein the finally, after_finish_remote_logs. A shutdown between them turns a real success into running then interrupted/unknown (provisioner.py:942). Commit status=complete with ready_at.
Prior art / DRY: reuses _manifest, attempt routes, _validated_hostname, require_admin_auth, and the existing dashboard conventions; reliability_snapshot duplicates tail()'s bounded source read (nit), and hardware_signature is a dead metadata write (nit).
Doc contradictions: origin-known policy is per-host in docs but global in code; invalid-start attempts are dropped only when a window is set; JS flattens 401/422 detail into a misleading "check your filters" message.
Credit: the CUDA probe compile grouping regression, the real setup-failure-to-launch-success series test, and the JS harness that executes the shipped file are all solid. Nits are optional; the three must-fix items should land before merge.
Track retention evictions per host so unrelated hosts retain measurable first attempts. Preserve and disclose uncertainty from older unattributed evictions. Convert local filter times to UTC for both reports and downloads. Commit successful status, readiness, and completion together. Preserve that outcome through final log retrieval and public provisioning cancellation while keeping unfinished shutdown attempts interrupted. Share bounded diagnostic reads, remove unused hardware metadata, expose API errors, restore dashboard page conventions, and document invalid-start filters. Cover timezone offsets, empty metrics, retention upgrades, source limits, and shutdown boundaries with regressions. Validation: 2,823 passed, 1 skipped; 93.79% branch-enabled coverage. Ruff lint and formatting, mypy, JavaScript syntax, and diff checks passed. The required retention, timezone, and shutdown regressions fail against the original PR code.
ttlogan
left a comment
There was a problem hiding this comment.
Re-read c888647 against the review findings. All of them land:
- Per-host evicted_attempts via the AFTER DELETE trigger, with unattributed_evicted_attempts carrying the older uncounted evictions, so one host's teardown no longer zeroes another host's first-attempt denominator.
- Filter bounds and the download both convert through new Date(value).toISOString(), so local input becomes real UTC and the exported query matches the displayed report.
- COMPLETE commits status, ready_at, and finished_at together; final retrieval and public cancellation preserve an already-complete outcome while interrupted shutdown attempts stay interrupted.
- harness_signature write is gone; the _source_records helper is shared by tail() and reliability_snapshot under one budget; the zero-denominator branch is asserted; API detail is surfaced; reliability.html matches the admin page conventions.
Metrics still hold (the L4 canary at 72.52 s under the 89.5 s threshold, 1/1 first-attempt success). Sound.
QIIP retains provisioning attempts and diagnostics but does not expose fleet success rates or recurring failures. This adds Admin > Fleet reliability, an authenticated reporting API, and JSON exports with UTC/cohort filters, explicit numerators and denominators, failure groups, and links to exact attempt logs and diagnostic bundles.
Attempts now retain retry-series identity across restarts and bundle changes, plus registered readiness. Reports preserve original errors beside versioned signatures and distinguish cancellation, explicit unsupported hardware, interrupted outcomes, missing environment evidence, legacy origins, and retention gaps. Reporting reads retained records without contacting nodes. Real fleet logs exposed two CUDA probe compilation failures split by surrounding output; these now group together without altering their original errors.
Closes #124. Real baseline and canary validation are complete, and the operator accepted the pilot targets on 2026-09-22: one successful first attempt, zero cancellations/unsupported outcomes, and readiness at or below 89.5 seconds (rounded from 120% of the observed baseline). The measured candidate meets every accepted target.
Real-fleet validation:
76d61ef: 1/1 first-attempt success, 72.521237-second readiness (one sample), zero cancellations and unsupported outcomes. Retry recovery remains unmeasured because no real provisioning failure occurred.READY. Full cycles including placement delays took 626.648662 and 636.394442 seconds respectively; these are distinct from setup readiness.docs/fleet-reliability.md.One sample per version establishes this pilot's outcome, not a fleet-wide performance improvement. Missing historical environment evidence and the unavailable-journal warning remain visible.
Automated validation:
76d61efpassed GitHub Quality and Python 3.13.The local suite used pinned uv 0.12.17 with the host's recursive
BASH_ENVhook removed. Browser visual review was unavailable; the live authenticated page/API and the JavaScript regression harness were checked.