Repository navigation
Conversation
A devnet paused for longer than its stability window (laptop lid closed, DevKit stopped for days) could never produce blocks again: a Haskell producer only forges while its tip is less than 3k/f slots behind wall clock (5 minutes on the default devnet). The only way out was reset, which throws away contracts, wallets and stake. DevKit now catches the chain up with Yano and keeps it (see ADR-0017): - Companion: the Haskell node restarts as a relay, Yano follows its chain, backfills sparse empty blocks to wall clock (one per forecast window plus every epoch boundary), the relay adopts them, and the node restarts as the producer. A 7-minute pause took 18 s. - Yano-only: start runs Yano without forging first, backfills, then starts the live producer, so every skipped epoch gets its boundary instead of one block jumping over them. It runs on start, from a watchdog every 30 s while a companion devnet runs (the closed-lid case), with the new catch-up command, and through POST /devnet/catch-up; GET /devnet/chain-lag reports the lag. devnet.auto.catch.up=false turns the automatic part off. Supporting changes: - YanoService takes a run mode (LIVE, PAST_TIME_TRAVEL, FOLLOW, CATCH_UP) and passes every property as an env var as well - cardano-node is stopped with SIGTERM to the node process, not a forced kill, so restarts do not revalidate the whole ChainDB (also for stop) - RelaySyncWaiter can wait for a slot, not only an epoch - warn when the installed Yano is not the pinned version: download keeps an existing binary, so a DevKit upgrade left pre6 in place Verified: lid-closed freeze of 6.5 min healed by the watchdog; stop, wait 7 min, start; yano-only stop across 4 epochs (boundaries 3->4 to 6->7 processed one by one); SDK e2e suites 13/13 on the caught-up chain. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A catch-up restarts the node twice and can take a while; the console used to
print a handful of ✅ lines with silences in between, plus leftover noise
("Waiting for node socket file ...", "Deleted pid file : yano.pid"), and a
watchdog catch-up was printed straight after the shell prompt.
New util.progress.ConsoleProgress renders work as numbered steps:
[1/4] ✅ Haskell node → relay ................. tip slot 1,067 (1s)
[2/4] ⠹ Yano following the chain ▕██████░░░░▏ 61% slot 92,310 / 150,004
a spinner while a step runs, a bar with a percentage when the total is known,
and each step's duration. In a terminal the line is redrawn in place;
anywhere else (Docker logs, pipes, the admin API's writer) the same steps are
plain lines with progress throttled to every 5 s. YACI_PROGRESS=plain forces
plain lines. Only ConsoleWriter.console() gets in-place rendering, so other
writers are never fed carriage returns.
Used for:
- devnet catch-up (4 steps; both syncs show real percentages)
- the companion first-run relay sync
- Yano and Yaci Store start-up (one spinner line instead of a line per
second)
- Yaci Store's sync to the chain tip (indexer height / node height)
- component downloads (MB / MB)
The watchdog now starts below the shell prompt and redraws it when done.
RelaySyncWaiter gains a numeric tip listener for the bars.
Verified in a pseudo-terminal: companion create, a 70 s freeze healed by the
watchdog with Ogmios, Kupo, submit-api and Yaci Store running (all
reconnected; SDK e2e suites 13/13 afterwards), and an Ogmios download.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
On a machine where the node socket took a while to appear, the relay sync step was interleaved with "Find tip error : Node Socket file is not available yet" on every poll, each glued to the end of the spinner line. - ClusterUtilService.getTip reports errors through the caller's writer instead of printing them directly. Polling callers (relay sync, catch-up, Yaci Store sync wait, admin API, MCP) already pass a silent or debug writer; the tip command still prints them. - While a progress line is shown in a terminal, System.out is wrapped: any other write clears the live line first and the spinner redraws it below, and output left without a trailing newline is not overwritten. Output from anywhere (relayed process stderr, warnings) now lands on its own line. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…d yano-only backfill - stop/reset cancels a running catch-up instead of waiting up to 3 minutes and then stopping next to it. The catch-up's thread is interrupted (its waits, Yano HTTP calls and sleeps all stop), it starts no further process, and its clean-up does not restart the Haskell node, so a stop cannot find the node resurrected and a reset cannot race a restart while deleting the database. - Yano-only: a failed backfill no longer starts live production when the gap crosses an epoch boundary (its first block would skip the missed epochs). Inside one epoch it still starts, since no boundary is skipped. - The catch-up command and POST /devnet/catch-up always backfill a yano-only devnet, whatever devnet.auto.catch.up says. With auto catch-up off, start now leaves a chain that is behind with Yano running but not forging (like a stalled companion producer) and points to 'catch-up', instead of starting live production over the gap. - Companion hand-back counts as done only after the producer restarted; a failed restart goes through the recovery path, which reports if it cannot recover either. DevnetCatchUpServiceTest covers each case. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…le Yano, fresh lag on fallback - A process is tracked as soon as it is spawned. If its start-up is interrupted (stop/reset cancelling a catch-up), times out or exits, it is terminated before the error propagates: Yano (YanoService registers it right after spawning and terminates it on any failed start, including the 30 s timeout that used to leave it running), the 1 s start check in ProcessUtil, and the node's socket wait. ProcessUtil.terminate stops a process tree gracefully then forcibly and also works on an interrupted thread. - New YanoRunMode.IDLE: Yano serves the chain and its HTTP API with no block producer and no client, so it can never forge. A yano-only start with auto catch-up off now waits in this mode instead of CATCH_UP, whose long block time only postponed forging (by an hour). CATCH_UP, which only lives for the duration of a catch-up, now uses Yano's maximum block time. - A failed yano-only backfill re-measures the lag before allowing live production inside one epoch: the failed call may have taken minutes. Tests: a real child process whose start is interrupted is gone afterwards; terminate on an interrupted thread; IDLE properties; the lag re-check with a clock that moves five epochs during the failing call. Also verified with the Yano binary: auto catch-up off, idle Yano does not forge for 20 s+, then 'catch-up' backfills 3 epochs and goes live. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Verdict: CHANGES REQUESTED Review of PR #196 at F1 [P1] Propagate catch-up cancellation to the enclosing companion startupLocation: Problem: Evidence: A temporary probe exercises the real Requested revision: Carry cancellation/fatal startup failure through to F2 [P1] Keep Yano non-forging when the auto-off tip check failsLocation: Problem: With Evidence: Two temporary probes separately return Requested revision: Allow LIVE only after a successful tip read and an explicit safe-lag decision. On an unknown tip, retain IDLE if possible or fail startup with a retry message. Cover readiness timeout and a failed second tip read. F3 [P1] Recheck wall-clock lag after a successful Yano-only backfillLocation: Problem: Every non-null catch-up response unconditionally allows LIVE. A long backfill can succeed at its original target while wall clock has advanced by several epochs during the work. Companion mode pumps catch-up until it is sufficiently close to current time; Yano-only mode does not. Starting LIVE at this point skips the newly elapsed epochs. Evidence: The pinned Yano Requested revision: Compare the returned tip with fresh wall clock after success and repeat catch-up until the remaining gap is safe for handoff, with a bounded failure path if it cannot keep up. Account for the stop/restart interval too. Add a successful-but-slow-response regression alongside the existing failed-call clock-advance test. Things to retain
Validation and scopeExecuted from an isolated checkout of the reviewed SHA, with ./gradlew test --tests '*ChainLagTest' --tests '*DevnetCatchUpServiceTest' --tests '*RelaySyncWaiterTest' --tests '*YanoRunModeConfigTest' --tests '*YanoServiceStartTest' --tests '*ConsoleProgressTest' --tests '*YanoInstalledVersionTest' -x walletUiBuild -x walletUiInstall -x generateGitProperties --console=plain
./gradlew -I /tmp/yaci-pr196-review.y7KV6L/probes.gradle test --tests '*YanoServiceBackfillIntervalTest' --tests '*YanoServiceSlotLeaderTimeTravelTest' reviewProbeTest -x walletUiBuild -x walletUiInstall -x generateGitProperties --console=plainResults: 58 existing tests passed; 0 failed, 0 errors, 0 skipped (48 in the first selection, 10 in the second). Four temporary probe tests passed, asserting the problematic observations described above. There is no separate The first test attempt failed in No implementation files were changed by this review. No approval is given at this SHA. Under the ADR process, any future design acceptance is not certification of implementation correctness or security. |
… safe lag or a fresh backfill - F1: a stop/reset during start tells the start to give up: it starts no further process (submit-api, producer restart) and reports failure, and the stop waits for the start before stopping what it started. - F2: with devnet.auto.catch.up=false, yano-only starts LIVE only after a tip read and a safe lag; an unreadable tip fails the start. A lag that would cross more than one epoch boundary also stays idle. - F3: the yano-only backfill repeats against fresh wall clock until the gap is within the handoff slack, with a bounded failure path. - catch-up runs for an idle yano-only Yano even before the lag reaches the stability window. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Author reply: round r1 · head 8f92d08 r1 is pushed as
Related fix found while doing F2 (not in the review): a yano-only devnet left IDLE with a lag between ¾ and 1 window could not be resumed: Deliberately unchanged: a companion restart whose catch-up fails for any reason other than a cancel still starts submit-api and reports success. The catch-up always leaves the Haskell producer running with the original topology, so the devnet is usable and Regression tests (new):
Validation (JDK 21.0.3-tem,
Over to you for round 2. |
Keep a devnet across long pauses. Close the laptop for the weekend or stop DevKit for days, and the same chain comes back: contracts, wallets, stake and governance state are kept, and no
resetis needed.Problem
A Haskell block producer can only forge while its tip is less than one stability window (
3k/fslots) behind wall clock: 5 minutes on the default devnet, 21 s with--epoch-length 40. After a longer pause the chain stalls for good, andresetwas the only way out:startwaits for a block that never comes, so Yaci Store, socat and Ogmios are not started either.What happens now
Companion: the Haskell node restarts as a relay, Yano follows its chain, then Yano backfills sparse empty blocks to wall clock. The backfill is one block per forecast window plus one at every epoch boundary, so rewards, snapshots and governance run for each epoch that passed. The relay adopts the blocks and restarts as the block producer.
Yano-only: on
start, Yano first runs without forging, backfills, then starts live.Triggers
startcatch-up [--force]--forcePOST /local-cluster/api/admin/devnet/catch-up?force=falseGET /local-cluster/api/admin/devnet/chain-lagstalled(for the viewer, MCP, scripts)devnet.auto.catch.up=falseturns off the automatic part:startand the watchdog then only print acatch-uphint.Changes
Catch-up (
localcluster/catchup/)ChainLag: lag, forecast window, thresholds, estimates.DevnetCatchUpService: companion and yano-only flows. Failures always leave the devnet with Yano stopped, the original topology and the Haskell producer running.DevnetStallWatchdog: armed and disarmed by theClusterStarted/Stopped/Deletedevents.stopwaits for a catch-up that is in progress.NodeControl, implemented byClusterStartService.Yano
YanoServicetakes a run mode (LIVE,PAST_TIME_TRAVEL,FOLLOW,CATCH_UP) and passes every property as an env var too.CATCH_UPuses a 1-hour block time, so no live block can land at the wall-clock slot before the backfill.downloadkeeps an existing binary, and a machine here was still on pre6.Node restarts: cardano-node is stopped with SIGTERM to the node process instead of a forced kill (also on
stop). A forced kill left the ChainDB unclean, and the next start revalidated every chunk: about 24 s, growing with the chain. A clean stop takes 0.1 s and the restart about 1.5 s.Progress UI (
util/progress/ConsoleProgress)YACI_PROGRESS=plainforces plain lines.Misc
RelaySyncWaitercan wait for a slot and report each tip read to a listener.getTipreports errors through the caller's writer, which stops the "Find tip error" spam from polling callers, the admin API and MCP.Docs: "Resuming After a Long Pause" in
yano-node-modes.mdx,catch-upincommands.mdx, and the property inapplication.properties.Verification
Default and short-epoch companion devnets and yano-only devnets, Yaci Store 2.0.6 native:
SIGSTOPped 6.5 min (default devnet) and 70 s (--epoch-length 100), with Store, submit-api, Ogmios and Kupo runningstart(default devnet)startcatch-upendpoint on a live chainscript)ChainLag, slot-targetRelaySyncWaiter, Yano run-mode config/env, installed Yano version, progress component)Limitations
activeSlotsCoeff1. Yano's catch-up on an existing chain does not check VRF eligibility; that needs a Yano change.yano-primaryandhaskell-onlyare not covered.Design notes: ADR-0017 (kept locally with the other ADRs).
🤖 Generated with Claude Code