Skip to content

feat: report provisioning reliability across the fleet - #153

Merged
sadsfae merged 6 commits into
quadsproject:developmentfrom
grafuls:feat/issue-124-fleet-reliability
Sep 22, 2026
Merged

sadsfae merged 6 commits into
quadsproject:developmentfrom
grafuls:feat/issue-124-fleet-reliability

Conversation

@grafuls

@grafuls grafuls commented Sep 21, 2026 •

Copy link
Copy Markdown
Contributor

QIIP retains provisioning attempts and diagnostics but does not expose fleet success rates or recurring failures. This adds Admin > Fleet reliability, an authenticated reporting API, and JSON exports with UTC/cohort filters, explicit numerators and denominators, failure groups, and links to exact attempt logs and diagnostic bundles.

Attempts now retain retry-series identity across restarts and bundle changes, plus registered readiness. Reports preserve original errors beside versioned signatures and distinguish cancellation, explicit unsupported hardware, interrupted outcomes, missing environment evidence, legacy origins, and retention gaps. Reporting reads retained records without contacting nodes. Real fleet logs exposed two CUDA probe compilation failures split by surrounding output; these now group together without altering their original errors.

Closes #124. Real baseline and canary validation are complete, and the operator accepted the pilot targets on 2026-09-22: one successful first attempt, zero cancellations/unsupported outcomes, and readiness at or below 89.5 seconds (rounded from 120% of the observed baseline). The measured candidate meets every accepted target.

Real-fleet validation:

  • Retained history: 10 provisioning attempts across five hosts, with 6 successes and 4 failures; all four retained failure bundles exported. Legacy retry origins and readiness remain unknown.
  • Matched baseline and candidate: the operator-selected NVIDIA L4, existing Qwen3.6-35B llama.cpp artifact and catalog profile, unchanged workload configuration. Each cycle gracefully tore down the node and let its existing placement claim reprovision the same profile.
  • Baseline: one successful attempt; observed setup-to-registration time 74.554863 seconds. Its legacy report correctly remains unmeasured for first-attempt and readiness metrics.
  • Candidate 76d61ef: 1/1 first-attempt success, 72.521237-second readiness (one sample), zero cancellations and unsupported outcomes. Retry recovery remains unmeasured because no real provisioning failure occurred.
  • Both restored workloads returned HTTP 200 and READY. Full cycles including placement delays took 626.648662 and 636.394442 seconds respectively; these are distinct from setup readiness.
  • The report, exact logs/bundle links, download, dashboard access, and authentication passed live checks. All recorded report fields except generation time survived a further gateway restart unchanged; all six original serving nodes remained healthy.
  • The candidate remains active on the development gateway in a separate release directory. The original checkout, local edits, configuration, data, and service-unit backup are preserved. Timestamped environment observations, attempt IDs, bundle versions, closed windows, and evidence hashes are in docs/fleet-reliability.md.

One sample per version establishes this pilot's outcome, not a fleet-wide performance improvement. Missing historical environment evidence and the unavailable-journal warning remain visible.

Automated validation:

  • Original full local suite: 2,811 passed, 1 skipped, with 93.72% branch-enabled coverage (92% required).
  • The CUDA grouping regression failed before the fix; all 20 reliability checks passed after it, including API, shipped setup/launch, and JavaScript checks. Reprocessing real history preserved every original error, outcome count, denominator, and evidence gap.
  • Deployed candidate 76d61ef passed GitHub Quality and Python 3.13.
  • Lint, formatting, type, script, and diff checks passed. Documentation follow-ups record the fleet evidence and accepted targets.

The local suite used pinned uv 0.12.17 with the host's recursive BASH_ENV hook removed. Browser visual review was unavailable; the live authenticated page/API and the JavaScript regression harness were checked.

@ttlogan ttlogan left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed the fleet reliability report. Each denominator keeps its numerator explicit (first-attempt success, retry recovery, cancellation, unsupported, readiness), and the docs match the code. Retry series identity surviving restarts and bundle changes is the right call, and old history aging into a conservatively unknown origin beats guessing. Ran the new suite: 19/19 pass, including the real setup-failure-to-launch-success path. CI is green. Looks good.

@grafuls
grafuls marked this pull request as ready for review September 21, 2026 21:07

@ttlogan ttlogan left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Follow-up on the cuda_probe_compile fix. The two probe failures that split by different log tails now land under one v1:cuda_probe_compile signature, and the regression test reproduces the split then confirms grouping. Ran it: passes. Docs baseline now records the real snapshot (10 attempts, 6/4, 9 llama.cpp / 1 vLLM) with the canary blocked on DNS, which is the honest state. Nothing to add.

@ttlogan ttlogan left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Read the matched L4 pilot record and re-ran the numbers. 120% of the 74.554863 s baseline is 89.465836 s, and the 72.521237 s canary lands at 97.27%, under the threshold. First-attempt and cancellation/unsupported denominators stay explicit, and the one-sample caveat is called out instead of over-claimed. Keeping the placement-inclusive recovery cycle (626/636 s) separate from the setup-to-readiness definition is right.

Looks good. No new findings on this docs commit.

@ttlogan ttlogan left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Re-read the accepted-targets doc change. The rounding is right: 120% of the 74.554863 s baseline is 89.465836 s, which rounds to the 89.5 s threshold, and the 72.521237 s canary still lands at 97.27%, under it. Status text now just records the operator acceptance instead of pending. No code touched, docs match the measured record. Looks good.

@sadsfae sadsfae left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review summary (multi-team pass)

The feature is well designed: stdlib-only report builder, no node I/O, explicit numerators/denominators, honest evidence gaps, and a real end-to-end recovered-series test plus a regression test for the CUDA probe grouping. Auth (require_admin_auth), route ordering, input validation, and the drill-down links all check out, and docs match the metric definitions. Three should-fix items remain.

Verification (head aa794ff): new reliability suites 20 passed (Python + shipped-JS harness); tests/provisioning 717 passed, 1 skipped, 1 env-only failure (test pins uv 0.12.17, host has 0.12.9); tests/api + tests/frontend 845 passed. Coverage per PR docs 93.72% branch; not re-measured here.

Must-fix

  1. Global evicted_attempts counter decides series_origin_known for every host, so any single eviction makes all later first-ever attempts unmeasured and the headline first_attempt_success metric silently degrades (log_store.py:168). Per-host eviction bookkeeping; the current test pins the global behavior.
  2. datetime-local values are browser-local but get :00Z, so non-UTC browsers filter and export a different window than labeled (reliability.js:78). Convert to ISO/UTC before sending.
  3. Success is committed in two writes: _update_state(COMPLETE) sets stage/ready_at but status stays running until _mark_log_complete in the finally, after _finish_remote_logs. A shutdown between them turns a real success into running then interrupted/unknown (provisioner.py:942). Commit status=complete with ready_at.

Prior art / DRY: reuses _manifest, attempt routes, _validated_hostname, require_admin_auth, and the existing dashboard conventions; reliability_snapshot duplicates tail()'s bounded source read (nit), and hardware_signature is a dead metadata write (nit).

Doc contradictions: origin-known policy is per-host in docs but global in code; invalid-start attempts are dropped only when a window is set; JS flattens 401/422 detail into a misleading "check your filters" message.

Credit: the CUDA probe compile grouping regression, the real setup-failure-to-launch-success series test, and the JS harness that executes the shipped file are all solid. Nits are optional; the three must-fix items should land before merge.

Comment thread inference_proxy/provisioning/log_store.py
Comment thread inference_proxy/static/js/reliability.js Outdated
Comment thread inference_proxy/provisioning/provisioner.py
Comment thread inference_proxy/provisioning/log_store.py Outdated
Comment thread inference_proxy/provisioning/reliability.py
Comment thread inference_proxy/static/js/reliability.js Outdated
Comment thread inference_proxy/templates/reliability.html
Comment thread tests/frontend/test_reliability_javascript.py
Comment thread inference_proxy/provisioning/log_store.py
Track retention evictions per host so unrelated hosts retain measurable first
attempts. Preserve and disclose uncertainty from older unattributed evictions.
Convert local filter times to UTC for both reports and downloads.

Commit successful status, readiness, and completion together. Preserve that
outcome through final log retrieval and public provisioning cancellation while
keeping unfinished shutdown attempts interrupted.

Share bounded diagnostic reads, remove unused hardware metadata, expose API
errors, restore dashboard page conventions, and document invalid-start filters.
Cover timezone offsets, empty metrics, retention upgrades, source limits, and
shutdown boundaries with regressions.

Validation: 2,823 passed, 1 skipped; 93.79% branch-enabled coverage. Ruff lint
and formatting, mypy, JavaScript syntax, and diff checks passed. The required
retention, timezone, and shutdown regressions fail against the original PR code.

@ttlogan ttlogan left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Re-read c888647 against the review findings. All of them land:

  • Per-host evicted_attempts via the AFTER DELETE trigger, with unattributed_evicted_attempts carrying the older uncounted evictions, so one host's teardown no longer zeroes another host's first-attempt denominator.
  • Filter bounds and the download both convert through new Date(value).toISOString(), so local input becomes real UTC and the exported query matches the displayed report.
  • COMPLETE commits status, ready_at, and finished_at together; final retrieval and public cancellation preserve an already-complete outcome while interrupted shutdown attempts stay interrupted.
  • harness_signature write is gone; the _source_records helper is shared by tail() and reliability_snapshot under one budget; the zero-denominator branch is asserted; API detail is surfaced; reliability.html matches the admin page conventions.

Metrics still hold (the L4 canary at 72.52 s under the 89.5 s threshold, 1/1 first-attempt success). Sound.

@sadsfae
sadsfae merged commit 1fc7b97 into quadsproject:development Sep 22, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants