Skip to content
This repository was archived by the owner on Sep 8, 2026. It is now read-only.
Merged
Show file tree
Hide file tree
Changes from 3 commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions crates/contributor-rewards/CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,8 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0

## [Unreleased]

- fix(contributor-rewards): verify the timestamp to Solana epoch mapping against real block times. `EpochFinder::find_epoch_at_timestamp` divided elapsed wall clock by a hardcoded 400ms slot duration and returned whatever epoch the result landed in, which drifts by roughly 30k slots per day of lookback on mainnet (about 7 percent of an epoch) and so picks the wrong epoch near a boundary. No fixed constant fixes that, because SIMD-0525 changes the real slot rate inside the lookback window. The estimate is now only a seed, and the epoch it lands in is confirmed against the first block of that epoch and of the next one, stepping one epoch at a time, tolerating skipped boundary slots, and rejecting a timestamp ahead of the chain tip rather than resolving it to whatever epoch a lagging endpoint sits in. The seed is set to 500ms, deliberately slower than any real cluster rate, so it lands at or after the epoch being looked for and the search only ever walks backward. A seed faster than the real rate would instead probe slots older than the target, which can fall outside the endpoint's retention and fail the search for a target that is itself readable. Because the epoch selects the leader schedule that contributor rewards are computed against, this now errors where it used to return a wrong answer: a backfill of a timestamp older than the RPC's ledger retention fails on the `ingestor::demand` path instead of silently mis-estimating. Epoch boundaries are resolved with one `getBlocksWithLimit` call rather than a capped forward walk of `getBlockTime` calls, so a long run of skipped slots at a boundary no longer fails permanently, and each search step reuses the bound the previous step resolved. `calculate_epoch_from_slot` is removed in favor of `EpochSchedule::get_epoch`, which is warmup-aware and does not underflow below `first_normal_slot` (malbeclabs/infra#2317)
- fix(contributor-rewards): `snapshot` validates before writing. It warns and continues when the leader schedule cannot be fetched, but `validate` treats a missing schedule as an error and every consumer rejects such a snapshot, so the command exited 0 having written an unusable file to the canonical `<network>-epoch-N-snapshot.json` name and a `snapshot` then `export-shapley` chain reported the first step green and failed on the second. The warn-and-continue predates this change; what changed is how reachable it is, now that resolving the Solana epoch depends on block-time reads rather than a single `getSlot` call. `snapshot` also no longer runs the epoch search twice per timestamp, taking the epoch from the leader schedule it already fetches (malbeclabs/infra#2317)
- migrate to Solana 3.0: workspace `solana-*` crates and `solana-sdk` move to the 3.0 line, `solana-program-test` to 3.0.12, and the doublezero SDK git-deps repin from `client/v0.27.1` to the malbeclabs/doublezero#3830 merge revision (malbeclabs/infra#1853)
- release artifact now builds as a static `x86_64-unknown-linux-musl` binary (malbeclabs/infra#1853)
- TLS for HTTP clients moves from openssl to rustls; trust roots are the bundled webpki Mozilla set plus the host OS certificate store, so OS-installed private CAs remain trusted (malbeclabs/infra#1853)
Expand Down
36 changes: 17 additions & 19 deletions crates/contributor-rewards/src/cli/snapshot.rs
Original file line number Diff line number Diff line change
Expand Up @@ -252,31 +252,22 @@ pub async fn create_snapshot(
fetcher.dz_rpc_client.clone(),
fetcher.solana_read_client.clone(),
);
let solana_epoch = match epoch_finder
.find_epoch_at_timestamp(fetch_data.start_us)
// fetch_leader_schedule resolves the Solana epoch itself and reports which
// one it used, so taking the epoch from its result avoids running the
// chain-verified epoch search twice over the same timestamp.
let leader_schedule = match epoch_finder
.fetch_leader_schedule(fetch_epoch, fetch_data.start_us)
.await
{
Ok(epoch) => Some(epoch),
Ok(schedule) => Some(schedule),
Err(e) => {
warn!("Failed to determine Solana epoch: {}", e);
warn!("Failed to get leader schedule: {}", e);
Comment thread
nikw9944 marked this conversation as resolved.
None
}
};

let leader_schedule = if solana_epoch.is_some() {
match epoch_finder
.fetch_leader_schedule(fetch_epoch, fetch_data.start_us)
.await
{
Ok(schedule) => Some(schedule),
Err(e) => {
warn!("Failed to get leader schedule: {}", e);
None
}
}
} else {
None
};
let solana_epoch = leader_schedule
.as_ref()
.map(|schedule| schedule.solana_epoch);

// Create metadata
let metadata = SnapshotMetadata {
Expand Down Expand Up @@ -327,6 +318,13 @@ pub async fn create_snapshot(
Network::Devnet => "dn",
};

// Refuse to write a snapshot no consumer can read. Every consumer rejects a
// snapshot with no leader schedule, and the write lands on the canonical
// `<network>-epoch-N-snapshot.json` name, so exiting 0 without this check
// leaves an unusable file where the next command expects a good one and
// reports the failure a step late.
snapshot.validate()?;

// Export: local override or configured storage
if local_file.is_some() || local_dir.is_some() {
// Save to local filesystem (ignores storage backend config)
Expand Down
Loading
Loading