Repository navigation
fix(beacon): persist unrealized justifications so get_head survives a restart - #628
MegaRedHand wants to merge 7 commits into
Conversation
… restart Beacon fork choice kept the per-block unrealized justified checkpoint (consensus-specs' store.unrealized_justifications[block_root]) only in an in-memory BeaconScratch map. A restart emptied that map with nothing to refill it: once a pre-restart block was a leaf from an epoch older than the store's clock, get_voting_source needed its entry and hit a hard SpecAssert on every tick, freezing get_head. Operators saw a follower that imported blocks and gossiped fine but never advanced its head after a restart or resume. Add a new storage table, BeaconUnrealizedJustifications, keyed by block root with (slot, Checkpoint) as the value. Root-keyed rather than slot||root like LiveChain/BlockProof: both readers of this table (get_voting_source, is_ffg_competitive) look up a root with no slot in hand, so a root-only key keeps that lookup a single point read; the slot rides in the value only for the pruner, which is fine since the table only ever holds the unfinalized window's worth of leaves between two finalizations. The in-memory map stays as a write-through cache in front of the table, since get_voting_source reads it for every block from a prior epoch. The write happens right after compute_pulled_up_tip's caller (on_block) has already committed the block's own insert_signed_block/insert_state, not in the same atomic batch: compute_pulled_up_tip needs the block's post-state, which insert_state hands off to the background state writer rather than committing synchronously. A crash in that narrow window leaves one block without its entry, surfacing as the same SpecAssert an unpatched build raised on every block, now scoped to a single block instead of the whole unfinalized window. Pruned on finalization inside update_checkpoints' Chain::Beacon arm, reusing the block_entry(&finalized.root) lookup PR #46 introduced for LiveChain pruning, on the same horizon (the finalized block's own slot rather than the stored checkpoint's epoch-start slot), so the finalized block's own entry always survives a resume with nothing imported past it yet. Bumps DB_VERSION to 4: a directory written by the previous version has no row in the new table for any block it already imported, which would otherwise look identical to "no unrealized justification computed yet" rather than "this directory predates the table".
…justifications # Conflicts: # CLAUDE.md # crates/storage/src/store.rs
…ng unrealized justification After the persistence fix, a missing unrealized justification is only possible in one narrow crash window: the state writer flushes a block's post-state, the process dies before that block's own entry lands, and resume's has_state check then skips re-importing it. Raising SpecAssert on that miss left get_voting_source (and, through it, filter_block_tree/ get_head) permanently unable to resolve a head through that leaf, for as long as it stayed a leaf. Prysm treats a rebuilt node's forkchoice store the same way: every block it seeds at startup gets the store's current justified checkpoint rather than a freshly computed unrealized one (buildForkchoiceChain, and the doubly-linked-tree store's insert). That value is always viable in filter_block_tree, since voting_source.epoch == store.justified.epoch holds by construction, so the fallback can only make a leaf more likely to survive filtering, never change what weight it casts once it does. get_voting_source now falls back to store.beacon_justified_checkpoint() on a miss instead of raising, logging at debug (get_head runs every tick, so a warn would repeat for as long as the block stays a leaf). The fallback is read-time only: it is never written into the table or the in-memory cache, since it is an approximation of the real value, and persisting it would later be read back as exact. is_ffg_competitive keeps raising on a miss; nothing in production calls it. Doc comments in tables.rs, store.rs and docs/data_storage.md that described a miss as fatal are updated to describe the crash window and the fallback. Store::delete_unrealized_justification is a new test-only helper that removes one entry from both the table and the cache, for simulating that crash window without a whole-process restart.
…justifications Picks up the preset-independent slots in the empty-boundary pruning test, which fixes both minimal-preset CI jobs here too.
🤖 Claude Code ReviewReview of PR 628: persist beacon unrealized justificationsThe approach is sound. It fixes a real Findings
Good
I only read the diff. I didn't build the code or run any tests. Automated review by Claude (Anthropic) · sonnet · custom prompt |
🤖 Kimi Code ReviewI'll review this PR which adds persistence for beacon unrealized justifications to fix a restart-induced SummaryThe PR correctly identifies and fixes a critical consensus bug: Issues Found1. Race Condition in Pruning: Cache Inconsistency Window (Critical)File: let mut batch = self.backend.begin_write().expect("write batch");
batch
.delete_batch(
Table::BeaconUnrealizedJustifications,
keys_to_delete.clone(),
)
.expect("delete pruned unrealized justifications");
batch.commit().expect("commit");
let mut scratch = self.beacon.lock().unwrap(); // Lock acquired AFTER commit
for key in &keys_to_delete {
scratch
.unrealized_justifications
.remove(&H256::from_slice(key));
}Problem: The in-memory cache is updated after the database commit. If a concurrent reader calls Fix: Acquire the let mut scratch = self.beacon.lock().unwrap(); // Lock FIRST
// ... collect keys_to_delete using scratch + backend view ...
let mut batch = self.backend.begin_write().expect("write batch");
batch.delete_batch(...).expect(...);
batch.commit().expect(...);
for key in &keys_to_delete {
scratch.unrealized_justifications.remove(&H256::from_slice(key));
}
// Lock drops here, after both structures are consistent2. Panic on Corrupted Value in
|
| Priority | Item | File |
|---|---|---|
| Critical | Fix cache-database ordering in prune | crates/storage/src/store.rs:3526 |
| High | Harden decode_unrealized_justification_value against corruption |
crates/storage/src/store.rs:744 |
| Medium | Add test verifying fallback doesn't incorrectly flip head | crates/blockchain/.../fork_choice.rs |
| Low | Document expect panic policy |
crates/storage/src/store.rs |
| Low | Rename/annotate import_unchecked |
crates/blockchain/.../fork_choice.rs |
Automated review by Kimi (Moonshot AI) · kimi-k2.6 · custom prompt
🤖 Codex Code ReviewLooks good overall. The restart-freeze root cause is addressed in the right place: persisting per-block unrealized justifications and pruning them on the same horizon as
No blocking issues from me. Automated review by OpenAI Codex · gpt-5.4 · custom prompt |
…realized-justifications
…i-626-63-64-633-636-638-gloas-live Both sides added storage tables and bumped DB_VERSION to 5: tmp for the gloas Config keys and its two gloas tables, #628 for BeaconUnrealizedJustifications. The tables are unioned (thirteen), and DB_VERSION goes to 6 with both reasons in its history, so a directory from either side's version 5 is refused rather than misread. Beyond the conflict markers: - get_forkchoice_store keeps tmp's anchor_slot, the anchor block's own slot, captured before the block moves. #628 shadowed it with the anchor state's slot, which a checkpoint-synced anchor advances past its block. That changed the slot tmp records with the anchor's payload link, and the anchor's unrealized-justification row is pruned against block slots, which on_block also records for every other row. - tmp's tests that call set_unrealized_justification pass the block's slot, which #628 added as a parameter. - BeaconScratch's doc and docs/data_storage.md list both sides' persisted exceptions, and drop unrealized_justifications from the uncapped maps, since #628 prunes it with its table.
Stacked on #627. It reuses #627's pruning cutoff, so it should merge after #627.
Why
Beacon fork choice keeps each block's unrealized justified checkpoint (consensus-specs'
store.unrealized_justifications[block_root]) only in memory, inBeaconScratch. A restart empties that map, and nothing refills it:on_blockby thehas_statechecks.Once a block imported before the restart is a leaf from an epoch older than the store clock,
get_voting_sourcefails withSpecAssert("block_root in store.unrealized_justifications").filter_block_treepasses that error up intoget_head.What an operator sees: no crash, but
warn "Failed to compute beacon head"on every tick, and the head frozen:lean_head_slotstays flat;forkchoiceUpdated;The head leaf from before the restart clears once a child of it is imported. A stale fork leaf keeps failing until justification moves past it.
Other clients
This PR does both: it persists the values the way Lighthouse does, and on a miss it falls back the way Prysm does (see "Fallback on a miss" below).
What
BeaconUnrealizedJustifications. The key is the block root; the value isslot (8 bytes, big-endian) ‖ Checkpoint SSZ.get_voting_source,is_ffg_competitive) look up by root without a slot. The slot is in the value only for the pruner.compute_pulled_up_tip(called fromon_block) and when the anchor is set up inget_forkchoice_store.Chain::Beaconarm ofStore::update_checkpoints, to fix(storage): prune the beacon LiveChain below the finalized block #627's cutoff: the finalized block's own slot. So the finalized block's own entry survives a resume that has imported nothing past it yet.DB_VERSION4 → 5. A data directory written by version 4 has no row for any block it already imported. (The first commit's message says 4; the merge frombeacon-chain-integration, which picked up lambdaclass/ethlambda_private#44's bump, makes it 5.)docs/data_storage.md(new table section, table diagram, pruning, what stays in memory) andCLAUDE.md(table count,DB_VERSIONbullet).Fallback on a miss (
5a7d6445)The entry is written right after the block's import, not in the same atomic batch, because
insert_statehands the post-state to the background state writer. So a crash can still leave one block with a state but no entry: the writer flushes the state, then the process dies before the entry is written. On restarthas_stateskips re-importing that block.For that case,
get_voting_sourceno longer raises on a miss. It returnsstore.beacon_justified_checkpoint(), the value Prysm gives every block it rebuilds at startup:filter_block_tree. It decides whether a leaf counts, never how much weight it casts, so the worst case is a stale fork leaf competing on weight, not a frozen or regressed head.debug!, notwarn!.get_headruns every tick, so a warning would repeat until the block stops being a leaf.is_ffg_competitivestill raises on a miss. Only the spec tests reach it.Store::delete_unrealized_justificationis a public test helper, following the existinginsert_live_chain_entry/delete_live_chain_entriesconvention, since another crate's tests call it.Tests
get_head_survives_a_restart_over_a_pre_restart_leafandget_head_survives_a_restart_past_a_pre_restart_fork_leafboth fail on the base branch and pass here. They import through the realcompute_pulled_up_tip, then reopen the store the waymain.rsresumes.an_unrealized_justification_persists_across_a_reopened_storereplaces the old test that pinned the value as in-memory only.A pruning test with an empty first slot of the finalized epoch: the finalized block's entry survives, and entries below it go.
get_voting_source_falls_back_to_the_justified_checkpoint_on_a_missandget_head_survives_a_dropped_unrealized_justification(the crash window: entry deleted from the table and the cache,get_headstill returns that block).After merging the base and adding the fallback:
ethlambda-storage --libethlambda-state-transition --lib beacon::fork_choiceethlambda-state-transition --lib beacon::gossipethlambda-blockchain --libbeacon_spec_tests -- fork_choice(mainnet fixtures)-D warningsDeploying
Needs a fresh checkpoint sync on every follower, since the build refuses a version-4 data directory. lambdaclass/ethlambda_private#44's bump to 4 needs one too, so deploying both together costs one checkpoint sync.
Moved
beacon-chain-integrationlives on this repo.beacon-chain-integration@c79fabd5, is merged into the branch.