Skip to content

SuspendLayerUtils: only suspend listed sub-root layers - #541

Open
bernardgut wants to merge 1 commit into
LINBIT:masterfrom
bernardgut:fix/suspend-subroot-filter-pr
Open

bernardgut wants to merge 1 commit into
LINBIT:masterfrom
bernardgut:fix/suspend-subroot-filter-pr

Conversation

@bernardgut

Copy link
Copy Markdown
Contributor

Fixes #540 (see also #524).

Problem

On DRBD ≥ 9.3.0, snapshots and clones stall when the Primary replicates to a diskful LUKS replica on
another node and a filesystem is mounted. That covers DRBD,LUKS,STORAGE with two diskful replicas, and
a diskless Primary with diskful LUKS replicas.

  • On the Primary, drbdadm suspend-io fails with "did not terminate within 5 seconds" (exit 20).
  • The satellite reports Failed to suspend IO … on layer DrbdLayer only when DRBD's ko-count drops
    the peer, about 42 s with the defaults.
  • Application writes on the Primary stay blocked until then.

Production evidence is in #540; the test-cluster reproduction is in the follow-up comment there.

Root cause

The dead filter. SuspendLayerUtils.setShouldSuspendStateRec
(controller/src/main/java/com/linbit/linstor/layer/utils/SuspendLayerUtils.java:70-86, identical on
master and v1.35.2) sets shouldSuspendIo on every layer it visits (line 77) and recurses into every
child (line 84).

  • The LAYERS_TO_SUSPEND_SUB_ROOT check at lines 80-83 therefore never takes effect: every layer is
    flagged, and listed children twice.
  • That contradicts the class and field Javadoc (lines 13-25).
  • It has been this way since 7047c32 (v1.33.0). Before that, only the root layer was flagged.

Why it started to matter. Until d5b4701 (v1.34.0) the extra flags did nothing.

  • LUKS volumes reported exists() == false, so SuspendManager skipped them (SuspendManager.java:205).
    STORAGE, NVME and BCACHE do not support suspend.
  • d5b4701 made LuksLayer set exists(true) (LuksLayer.java:405).
  • Since then the non-root pass (SuspendManager.java:105-121) runs dmsetup suspend on the dm-crypt
    device below DRBD on every node whose own drbdadm suspend-io succeeded.

How the stall happens. DRBD ≥ 9.3.0 suspend-io calls bdev_freeze() first, then sets
susp_user, then waits for the device to drain (drbd_nl.c:6608-6651 @ drbd-9.3.4; LINBIT/drbd@27ca01bf67).
Together:

  1. On the Primary, the XFS freeze writes at least a superblock log record and then the superblock.
    Under protocol C each completes only after the peer's P_WRITE_ACK.
  2. A peer with nothing mounted returns from its DRBD suspend at once. About 0.1–0.4 s after the
    request it suspends its dm-crypt device.
  3. A freeze write that reaches the peer after that gets no ack until the dm-crypt device is resumed.
    The Primary's drbdsetup stays in bdev_freeze until ko-count or an external resume.
  4. drbdadm's 5 s timeout makes drbdadm exit, but drbdsetup holds the inherited pipes, so
    ExtCmd.syncProcess() keeps the Primary's DeviceManager blocked.
  5. The controller only aborts and resumes after every satellite has answered, so nothing on the
    LINSTOR side ends the wait.

Change

What changes. Flag the topmost layer unconditionally, as documented, and lower layers only if they
are in the given set. The recursion still descends through unlisted layers. Resume still clears every
layer (ALL_LAYERS), now exactly once.

Scope of the change. Only the controller changes; the satellite and the wire format are untouched.
This was validated with stock 1.35.2 satellites and a 1.35.2 controller built from the same change.

Difference from the #540 sketch. The sketch used getParent() == null. Here
setShouldSuspendStateRec flags the root, and setShouldSuspendStateOfChildrenRec walks the children
and keeps the existing per-child filter. The behaviour is the same, and it does not depend on parent
pointers.

Why snapshots stay consistent

The controller takes the storage snapshot only after every node has answered the suspend update
(including its SuspendManager run), and only after all DRBD volumes are UpToDate
(CtrlSnapshotCrtApiCallHandler.java:536-547, 744-768).

On the Primary. A suspend-io that drbdadm reports as successful has completed bdev_freeze() and
the io_drained() wait: local_cnt == 0 and ap_pending_cnt == 0 for every peer
(drbd_nl.c:6588-6651). Any drbdadm failure aborts the snapshot.

  • Under protocol C, a write stays in ap_pending_cnt until the peer's P_WRITE_ACK arrives.
  • A peer sends P_WRITE_ACK only after the write completed on its backing device (e_end_block_tail,
    drbd_receiver.c:3161-3186 @ drbd-9.3.4).
  • dm-crypt completes a write only after the encrypted clone completed below it (crypt_endio →
    crypt_dec_pending → bio_endio(base_bio), drivers/md/dm-crypt.c:1819-1848, 1867-1896 @ v6.18).

So once the Primary's suspend has succeeded, every write the Primary issued is in each diskful
replica's STORAGE volume, below dm-crypt, which is what gets snapshotted.

On a peer. susp_user does not block replication. It only gates local application IO
(may_inc_ap_bio, drbd_int.h:3200-3205), and receive_Data has no suspend check
(drbd_receiver.c:4121-4329). What blocked the acks was suspending the dm-crypt device below DRBD.

Why this is enough.

  • This gives DRBD,LUKS,STORAGE the same guarantee as DRBD,STORAGE, and the behaviour it had up to 1.33.
  • dm-crypt has no cache to flush, unlike WRITECACHE and CACHE, which LAYERS_TO_SUSPEND_SUB_ROOT exists for.
  • DRBD's freeze stays: it is what makes the snapshotted filesystem clean, which readers mounting a
    snapshot norecovery depend on.
  • A cross-node barrier in SuspendManager would also avoid the stall, but it is a larger change than
    restoring the documented filter.
  • bab0e23, e10406c and b0a45cf remain needed: for LUKS-as-root, for CACHE, and for LUKS devices
    that a stock controller left suspended. DRBD-rooted stacks simply no longer suspend LUKS.

Behaviour per stack

Stack Flagged before → after Satellite effect
DRBD,LUKS,STORAGE DRBD, LUKS, STORAGE → DRBD dm-crypt below DRBD no longer suspended
LUKS,STORAGE LUKS, STORAGE → LUKS none (topmost LUKS still suspended)
DRBD,WRITECACHE,STORAGE all → DRBD, WRITECACHE none (WRITECACHE below the root only flushes, WritecacheLayer.java:165-186)
DRBD,CACHE,STORAGE all → DRBD, CACHE none (CACHE below the root: cleaner flush, then dmsetup suspend, CacheLayer.java:180-250)
DRBD,STORAGE; NVME/BCACHE/STORAGE below the root all → root (+ listed) none (isSuspendIoSupported() is false for NVME, BCACHE, STORAGE)
WRITECACHE or CACHE root with LUKS below all → root (+ listed) LUKS no longer suspended

I also checked this mechanically, running stock and patched SuspendLayerUtils over all 216 linear
stacks that LayerUtils.isLayerKindStackAllowed accepts. Results:

  • Flag sets differ in 215 of them.
  • Effective satellite suspend calls differ in 83. "Effective" means: flagged, with
    isSuspendIoSupported(), and for non-root layers a root that was suspended in the same run.
  • All 83 differences are the same one: LUKS below a DRBD, WRITECACHE or CACHE root is no longer
    suspended.
  • Clones use the same helper for the clone source (CtrlRscDfnApiCallHandler.java:1125 on master),
    so they change the same way.

Tests

New src/test/java/com/linbit/linstor/layer/utils/SuspendLayerUtilsTest.java runs the real recursion
on mocked layer trees. All stacks are allowed by LayerUtils. Flagged layers are checked with
times(1), the rest with never().

Test Expected flags
DRBD,LUKS,STORAGE DRBD only
LUKS,STORAGE LUKS only
DRBD,WRITECACHE,STORAGE (data and cache STORAGE children) DRBD and WRITECACHE
DRBD,CACHE,STORAGE DRBD and CACHE
DRBD,BCACHE,WRITECACHE,STORAGE DRBD and WRITECACHE (descends through the unlisted BCACHE)
DRBD,NVME,LUKS,STORAGE DRBD only
resume on DRBD,LUKS,STORAGE every layer cleared exactly once

All 7 fail against master's SuspendLayerUtils. The WRITECACHE, CACHE and resume tests fail only
because stock sets listed lower layers twice.

Local runs on master 87450b1 with JDK 21:

Check Result
SuspendLayerUtilsTest 7/7
RscDfnCloneApiTest / SnapshotApiTest / SnapshotRestoreApiTest / SnapshotRollbackApiTest / LayerResourceIdDbDriverTest 25/25, 31/31, 17/17, 10/10, 5/5
:controller:checkstyleMain no findings in the changed file
ErrorProne -Werror compile of :controller OK
Full suite left to CI

Validation on a test cluster

Same setup as the reproduction on #540:

  • Talos 1.14.1 (kernel 6.18.51, DRBD 9.3.3), piraeus-operator 2.12.0.
  • DRBD,LUKS,STORAGE on LVM thin, two diskful replicas plus one diskless, protocol C.
  • XFS on the Primary; no external resume.

"This PR" means stock 1.35.2 satellites with a 1.35.2 controller built from the same change on v1.35.2
(the SuspendLayerUtils diff is identical).

stock 1.35.2 this PR
sequential snapshots (writer running) 9 attempts: 7 hung until ko-count (42.3–43.0 s), 1 Ready, 1 aborted (peer not UpToDate after a ko-count disconnect) 50/50 Ready in 2.5–2.6 s, each after a write + sync
layer_resource_suspended=true rows DRBD, LUKS, STORAGE on both diskful nodes; DRBD, STORAGE on the tiebreaker DRBD only; 0 LUKS, 0 STORAGE
Linstor-Crypt-* seen suspended (0.2 s sampling) the peer's, in every hung attempt none on DRBD-rooted volumes (only the LUKS-as-root test volume, as expected)
concurrent snapshots, 8 busy volumes, waves of 8/24/48 not run on 1.35.2 8/8 (30 s), 24/24 (60 s), 48/48 (100 s); no faulty resources, nothing left suspended or behind
restore: read-only norecovery mount, marker checksum – OK
single-diskful DRBD,LUKS,STORAGE; LUKS,STORAGE; clone (checksum); online resize; nothing left suspended – 3/3; 3/3; OK; OK; OK
Failed to suspend IO reports / ko-count disconnects / ha-controller evictions 7 / 7 / 3 0 / 0 / 0

Rebase and backport

Written on v1.35.2, which we run, then rebased onto master.

  • Backport: the code change applies unchanged to v1.35.2, and that build is what was validated.
    We'd appreciate it in a 1.35.x patch release.
  • During the rebase: the test was extended, and the CHANGELOG entry moved into the existing
    [Unreleased] → ### Fixed list. That hunk needs adjusting for a 1.35.x backport.
  • Interfaces: no REST, client or documentation changes.

Out of scope (possible follow-ups)

Unbounded pipe-EOF wait. ExtCmd.syncProcess() waits for EOF on the child's pipes without a bound
after the child exits (ExtCmd.java:178-180, OutputReceiver.java:210-226). Here drbdsetup held
them for about 37 s after drbdadm had exited.

  • A bounded wait alone is not enough. The stuck suspend-io holds the resource's adm_mutex, so a
    resume-io issued meanwhile can give up after drbdadm's 5 s.
  • The suspend then completes with nothing left to resume it, so a later re-check of suspended:user
    would also be needed.

No suspend-phase timeout. The suspend phase of snapshot creation has none; only the take-snapshot
step does (CtrlSnapshotCrtApiCallHandler.java:795-799).

Freeze longer than 5 s. drbdadm's 5 s cmd-timeout-short still bounds suspend-io, so a freeze
that must write back a lot of dirty data can fail at about 5–10 s, with no dm-crypt involved. This
predates the change.

vgs hang (LVM). The LVM VG existence check runs vgs with an empty VG set, so it gets no
--config and no ignore_suspended_devices=1 (LvmUtils.checkVgExistsBoolImpl →
getVgsInfo(emptySet), LvmCommands.java:89-93). b0a45cf covered only pvdisplay and vgscan.

  • On the LVM bench this hung a free-space query on the peer for about 2 min and delayed that node's
    resume.
  • vgs survived SIGKILL; most likely it was blocked reading the suspended dm-crypt device, but no
    stack was captured.
  • I will file it separately.

DRBD,CACHE,STORAGE. CACHE below DRBD is still dmsetup suspended on every node whose DRBD suspend
succeeded, so with DRBD ≥ 9.3 it may hit the same ordering problem. Not tested and not changed here.

Note that this PR was created with the help of Claude-Opus-5.5 and might contain some hallucinations/inconsistancies as I am not that familiar with the source code. Take everything it says on "root cause" onwards with a grain of salt. Thanks

setShouldSuspendStateRec() set shouldSuspendIo on every layer object it
visited and recursed into every child, so every layer of a resource was
flagged (listed children even twice) and the LAYERS_TO_SUSPEND_SUB_ROOT
filter (WRITECACHE, CACHE, DRBD) never took effect. This has been the
case since SuspendLayerUtils was introduced in 7047c32 / v1.33.0;
before, only the root layer was flagged.

The satellite only suspends a flagged layer if its volumes exist. Since
d5b4701 / v1.34.0 the LUKS layer reports its volumes as existing, so
snapshots and clones also make the satellite run "dmsetup suspend" on
the dm-crypt device of a LUKS layer below DRBD, on every node whose own
DRBD suspend succeeded. The suspended LUKS device behind the resume
hang fixed in e10406c has the same cause.

With DRBD >= 9.3, where an admin suspend-io freezes the mounted
filesystem, this stalls snapshots whenever the Primary replicates to a
diskful LUKS replica on another node (two diskful replicas, or a
diskless Primary): the freeze writes on the Primary need the protocol C
ack of that peer, and once the peer has suspended its dm-crypt device
below DRBD, drbdsetup suspend-io on the Primary blocks until ko-count
ejects the peer or the dm-crypt device is resumed.

Flag the topmost layer unconditionally, as documented, and lower layers
only if their kind is in the given set, while still descending through
unlisted layers. Resume still clears every layer, now exactly once.

Fixes: LINBIT#540
Authored-By: Bernard Gütermann <bernard.gutermann@sekops.ch>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Snapshot of DRBD,LUKS,STORAGE with two diskful replicas deadlocks: SuspendLayerUtils also suspends the LUKS layer (sub-root filter is dead code)

1 participant