feat: client-side compression/dedup via VDO (issue #277) - #402
Draft
boddumanohar wants to merge 29 commits into
Draft
boddumanohar wants to merge 29 commits into
boddumanohar wants to merge 29 commits into
Conversation
boddumanohar
added a commit
that referenced
this pull request
Aug 7, 2026
Tested against the real implementation (PR #402), not just raw LVM commands: two real PVCs with clientCompression/clientDeduplication on the same node, distinct checksummed data, node rebooted. Both VDO instances reattached cleanly via fresh NodeStageVolume calls (kubelet's own bookkeeping resets on reboot too) -- kvdo module usage count exactly 2, both VDOOperatingMode normal, both checksums matched exactly. Caveat found and documented: the two NodeStageVolume LVM command sequences happened to complete sequentially rather than genuinely overlapping, so LVM's internal command locking under truly concurrent vgchange/pvscan calls remains unexercised -- narrowed the open item accordingly rather than closing it outright. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
boddumanohar
added a commit
that referenced
this pull request
Aug 7, 2026
…ation Tested both clone paths against PR #402's real ResolveClonedVDO: a direct PVC-to-PVC clone and a snapshot restore, both scheduled onto the same node as their still-live source (the specific co-location scenario this finding warns about). Both correctly resolved via vgimportclone + lvrename, mounted cleanly with data matching the source exactly, and coexisted with the source and each other with independent VG identities and no cross-contamination. Also corrected the "Detection" section: the implementation ended up simpler than originally planned -- detection is unconditional and purely device-identity-based, not gated on VolumeContentSource, so no separate content-source plumbing was needed. Found and fixed one bug along the way: the collision-detection log message was picking up pvs's stderr WARNING: lines merged into its output instead of just the VG name -- harmless in this run, but fixed to parse cleanly. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
boddumanohar
added a commit
that referenced
this pull request
Aug 7, 2026
Deliberately reproduced against PR #402's real implementation: forcibly disconnected a VDO volume's NVMe-oF subsystem at the host level while the node stayed up, then deleted the pod. This exposed two real bugs (both now fixed on that branch): DeactivateVDO had no fallback for an unreachable device, and once added, the fallback's device-name matching didn't account for device-mapper's dash-escaping and matched nothing. With both fixed, cleanup is now fully automatic -- confirmed by reproducing the whole sequence a second time. Also documented an unplanned but valuable side observation: for ~19s after disconnect, cached reads/writes silently appeared to succeed before the real I/O failure surfaced, at which point VDO correctly fenced itself into read-only mode and ext4 independently aborted its journal -- both layers protected data correctly with no wiring needed from this design. Narrowed the remaining open item: the node-NotReady-and-rejoin path is still unverified (this test kept the node itself healthy throughout, only the storage connection was severed). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
3 tasks
boddumanohar
added a commit
that referenced
this pull request
Aug 11, 2026
…lap as a known gap #403 doesn't actually block #402 (it says so explicitly, and #402 already works around it with a nodeSelector pin for its own verification), so there's no need to carry the operator-side plumbing this fix grew to close a real but narrow edge case: a single Kubernetes node listed in more than one DHCHAP-gated pool's AllowedNodes. Revert the dhchap_node_label StorageClass parameter and the operator change that set it. hardPinTopologySegments goes back to matching the "simplyblock.io/pool." prefix, with the multi-pool-overlap behavior now documented as a deliberately accepted limitation (in both a code comment and a test that documents current behavior rather than asserting correctness) instead of solved. operator/ is now byte-for-byte identical to origin/main; this PR is CSI-driver-only again. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
boddumanohar
added a commit
that referenced
this pull request
Aug 12, 2026
#403) CreateVolume never populated AccessibleTopology for StorageClasses that select their cluster directly via cluster_id (the common case), so external-provisioner never set PV.spec.nodeAffinity. A bound PVC could then be rescheduled onto any node, even when the StorageClass gates provisioning to specific nodes via AllowedTopologies (DHCHAP's allowed-node label today, and VDO's node-local state in the upcoming #402). Extract only the segments that represent a genuine per-node constraint from the CSI AccessibilityRequirements and echo them back, leaving plain NVMe-oF-backed volumes (which behave identically from any node) unpinned. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
boddumanohar
added a commit
that referenced
this pull request
Aug 12, 2026
…lap as a known gap #403 doesn't actually block #402 (it says so explicitly, and #402 already works around it with a nodeSelector pin for its own verification), so there's no need to carry the operator-side plumbing this fix grew to close a real but narrow edge case: a single Kubernetes node listed in more than one DHCHAP-gated pool's AllowedNodes. Revert the dhchap_node_label StorageClass parameter and the operator change that set it. hardPinTopologySegments goes back to matching the "simplyblock.io/pool." prefix, with the multi-pool-overlap behavior now documented as a deliberately accepted limitation (in both a code comment and a test that documents current behavior rather than asserting correctness) instead of solved. operator/ is now byte-for-byte identical to origin/main; this PR is CSI-driver-only again. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
boddumanohar
added a commit
that referenced
this pull request
Aug 21, 2026
Tested against the real implementation (PR #402), not just raw LVM commands: two real PVCs with clientCompression/clientDeduplication on the same node, distinct checksummed data, node rebooted. Both VDO instances reattached cleanly via fresh NodeStageVolume calls (kubelet's own bookkeeping resets on reboot too) -- kvdo module usage count exactly 2, both VDOOperatingMode normal, both checksums matched exactly. Caveat found and documented: the two NodeStageVolume LVM command sequences happened to complete sequentially rather than genuinely overlapping, so LVM's internal command locking under truly concurrent vgchange/pvscan calls remains unexercised -- narrowed the open item accordingly rather than closing it outright. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
boddumanohar
added a commit
that referenced
this pull request
Aug 21, 2026
…ation Tested both clone paths against PR #402's real ResolveClonedVDO: a direct PVC-to-PVC clone and a snapshot restore, both scheduled onto the same node as their still-live source (the specific co-location scenario this finding warns about). Both correctly resolved via vgimportclone + lvrename, mounted cleanly with data matching the source exactly, and coexisted with the source and each other with independent VG identities and no cross-contamination. Also corrected the "Detection" section: the implementation ended up simpler than originally planned -- detection is unconditional and purely device-identity-based, not gated on VolumeContentSource, so no separate content-source plumbing was needed. Found and fixed one bug along the way: the collision-detection log message was picking up pvs's stderr WARNING: lines merged into its output instead of just the VG name -- harmless in this run, but fixed to parse cleanly. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
boddumanohar
added a commit
that referenced
this pull request
Aug 21, 2026
Deliberately reproduced against PR #402's real implementation: forcibly disconnected a VDO volume's NVMe-oF subsystem at the host level while the node stayed up, then deleted the pod. This exposed two real bugs (both now fixed on that branch): DeactivateVDO had no fallback for an unreachable device, and once added, the fallback's device-name matching didn't account for device-mapper's dash-escaping and matched nothing. With both fixed, cleanup is now fully automatic -- confirmed by reproducing the whole sequence a second time. Also documented an unplanned but valuable side observation: for ~19s after disconnect, cached reads/writes silently appeared to succeed before the real I/O failure surfaced, at which point VDO correctly fenced itself into read-only mode and ext4 independently aborted its journal -- both layers protected data correctly with no wiring needed from this design. Narrowed the remaining open item: the node-NotReady-and-rejoin path is still unverified (this test kept the node itself healthy throughout, only the storage connection was severed). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
boddumanohar
force-pushed
the
issue-277-client-side-compression-impl
branch
from
August 21, 2026 12:25
88d5be8 to
d201b02
Compare
boddumanohar
added a commit
that referenced
this pull request
Aug 25, 2026
Tested against the real implementation (PR #402), not just raw LVM commands: two real PVCs with clientCompression/clientDeduplication on the same node, distinct checksummed data, node rebooted. Both VDO instances reattached cleanly via fresh NodeStageVolume calls (kubelet's own bookkeeping resets on reboot too) -- kvdo module usage count exactly 2, both VDOOperatingMode normal, both checksums matched exactly. Caveat found and documented: the two NodeStageVolume LVM command sequences happened to complete sequentially rather than genuinely overlapping, so LVM's internal command locking under truly concurrent vgchange/pvscan calls remains unexercised -- narrowed the open item accordingly rather than closing it outright. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
boddumanohar
added a commit
that referenced
this pull request
Aug 25, 2026
…ation Tested both clone paths against PR #402's real ResolveClonedVDO: a direct PVC-to-PVC clone and a snapshot restore, both scheduled onto the same node as their still-live source (the specific co-location scenario this finding warns about). Both correctly resolved via vgimportclone + lvrename, mounted cleanly with data matching the source exactly, and coexisted with the source and each other with independent VG identities and no cross-contamination. Also corrected the "Detection" section: the implementation ended up simpler than originally planned -- detection is unconditional and purely device-identity-based, not gated on VolumeContentSource, so no separate content-source plumbing was needed. Found and fixed one bug along the way: the collision-detection log message was picking up pvs's stderr WARNING: lines merged into its output instead of just the VG name -- harmless in this run, but fixed to parse cleanly. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
boddumanohar
added a commit
that referenced
this pull request
Aug 25, 2026
Deliberately reproduced against PR #402's real implementation: forcibly disconnected a VDO volume's NVMe-oF subsystem at the host level while the node stayed up, then deleted the pod. This exposed two real bugs (both now fixed on that branch): DeactivateVDO had no fallback for an unreachable device, and once added, the fallback's device-name matching didn't account for device-mapper's dash-escaping and matched nothing. With both fixed, cleanup is now fully automatic -- confirmed by reproducing the whole sequence a second time. Also documented an unplanned but valuable side observation: for ~19s after disconnect, cached reads/writes silently appeared to succeed before the real I/O failure surfaced, at which point VDO correctly fenced itself into read-only mode and ext4 independently aborted its journal -- both layers protected data correctly with no wiring needed from this design. Narrowed the remaining open item: the node-NotReady-and-rejoin path is still unverified (this test kept the node itself healthy throughout, only the storage connection was severed). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Implements the design in design-issue-277-client-side-compression.md (PR #398): new clientCompression/clientDeduplication Pool StorageClass params, VDO-capable node topology gating, a new csi-driver/pkg/util/vdo.go managing per-volume VDO stacks over LVM, and nodeserver.go wiring to create/reattach/grow/remove VDO devices across stage/unstage/restage/expand. Not yet covered (tracked as follow-ups): unit tests, deliberate exercise of clone/snapshot VDO resolution, multi-instance+reboot, XFS-on-VDO, and crash-consistency of async write policy. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
goconst nudged a look at nodeserver.go's raw == "true" checks against client_compression/client_deduplication: mergeStorageClassParameters actually emits "True" (capitalized), so those checks would never have matched. Replaced with kube.BoolParam (the same parser already used for encryption/replicate), factored into a shared vdoParams helper. Also regenerated dist/install.yaml (missed by the earlier make manifests run) and fixed a goconst/unparam/lll nit in vdo.go and simplyblockpool_controller.go. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Found live on the test cluster: CSINode's topology key set is captured once
at CSI plugin registration (seconds after pod start), while
advertiseVDOCapability patches the underlying node label asynchronously in
the background (can take minutes, gated on the postStart hook's dnf
install). Conditionally omitting the vdo-capable key when false meant it was
essentially never present at registration time, permanently breaking the
topology gate for that node until its csi-node pod restarted -- confirmed by
a real ProvisioningFailed error ("topology ... is not in requisite") when
provisioning against a Pool with clientCompression/clientDeduplication
enabled. Now always present, matching the existing
topologyKeyStorageNodeUUIDPrefix pattern's own documented rationale: only
the key's presence needs to be stable, not its value, since
external-provisioner reads the live Node label value fresh on every
provision.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
pkg/util/vdo.go execs pvcreate/vgcreate/lvcreate/vgchange/lvextend/dmsetup etc. from inside the csi-node container, but base_image only ships nvme-cli/e2fsprogs/xfsprogs -- confirmed live on the test cluster (MountVolume.MountDevice failing with "pvcreate: executable file not found in $PATH"). Added to this branch-tagged Dockerfile rather than base_image/Dockerfile_base, since that image tag is shared across every branch and rebuilding it would affect unrelated in-flight work; worth promoting there once this feature lands. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…iner) Confirmed live: lvcreate --type vdo failed with "device not cleared, Aborting. Failed to wipe start of new LV" -- device-mapper's default behavior waits on udev to create/settle the resulting device node, but this container has no udev daemon running. DM_DISABLE_UDEV=1 is the standard fix for LVM tooling run inside a container. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
lvm2 alone wasn't enough: lvcreate --type vdo shells out to vdoformat
internally to format the new VDO pool, and that binary ships in the separate
vdo package, not lvm2. Confirmed live ("/usr/bin/vdoformat: execvp failed:
No such file or directory").
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
The multi-arch build was failing entirely on the arm64 leg: "No match for argument: vdo" -- Oracle Linux 9's configured repos don't carry a vdo build for aarch64. Client-side VDO is x86_64-only for now anyway (matches every host used in this design's validation), so gate the install on TARGETARCH rather than block the whole multi-arch image. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Critical bug caught live: deleting and recreating a pod on the same node (same PVC) came back with an empty filesystem -- all test data gone. NodeUnstageVolume was calling RemoveVDO (vgchange -an + vgremove -f) on every routine unstage, but NodeUnstageVolume fires any time no pod on this node currently needs the volume mounted, not only when the volume is actually being deleted. vgremove destroys the VG's LVM metadata, which is what makes VDO's compressed/deduplicated physical layout decodable back into the original file bytes -- so this was silently destroying user data on an entirely ordinary pod restart, directly contradicting the design's own "must be re-included every time the node-side CSI driver re-provisions the volume" requirement. Added DeactivateVDO (vgchange -an only, non-destructive, reversible via CreateOrAttachVDO's existing vgchange -ay reactivate path) and wired NodeUnstageVolume to use it instead. RemoveVDO is kept for genuine destroy/cleanup scenarios, just no longer called from the routine unstage path. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Confirmed live: lvextend -l100%FREE failed with "New size given (1024 extents) not larger than existing size (1535 extents)" during a real PVC expand (6Gi -> 10Gi). Unlike lvcreate, lvextend's bare "100%FREE" is an absolute target (100% of currently-free space alone), not "grow by" -- after the backend resize, free space (1024 extents) was smaller than the pool's current size (1535 extents), so the absolute interpretation rejected it as not larger. The "+" prefix makes it additive (current size + free space), which is the actual intent. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Confirmed live during a real clone test: ResolveClonedVDO's own log message
came out polluted with LVM's duplicate-PV warnings ("WARNING: Not using
device /dev/nvme2n1 for PV ...") ahead of the actual VG name, because
runLVMCommand merges stdout+stderr and pvVGName trusted the whole trimmed
blob. Didn't cause a functional problem in that run (the dirty string still
correctly differed from the target VG name), but is a latent correctness
risk and produced a confusing log line. Now takes the first non-empty,
non-WARNING line instead.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Critical gap confirmed live: forcibly disconnected a VDO volume's NVMe-oF connection at the host level (storage-side disconnect while the node stays up), then deleted the pod. NodeUnstageVolume's DeactivateVDO call failed with "Volume group ... not found" on every one of 18 retries -- the exact same failure mode vgremove hits in RemoveVDO when the backing device is gone, since vgchange -an also needs to read/write VG metadata that no longer exists. Kubelet eventually force-removed the pod anyway, leaving the orphaned dm-vdo stack permanently stuck with nothing left to clean it up -- my earlier fix (switching NodeUnstageVolume from the destructive RemoveVDO to the safe DeactivateVDO) accidentally dropped the dmsetup fallback robustness RemoveVDO already had for this exact case. DeactivateVDO now falls back to the same direct dmsetup removal when vgchange -an fails with a "not found"-style error (mirroring vgExists' existing pattern-matching), gated on that specific failure signature so a genuinely busy/in-use VG (device still reachable) is never forced. Kernel logs from the same test also confirmed VDO's own fencing behavior worked correctly and safely: once the disconnected device's I/O actually started failing (~19s after disconnect, buffered writes had masked it until then), VDO fenced itself into read-only mode and ext4 aborted its journal, rather than corrupting data. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Confirmed live: the dmsetup fallback added for DeactivateVDO correctly
triggered ("vgchange -an failed ... falling back"), but matched zero dm
device names and left the orphaned stack in place. device-mapper flattens
"<vg>-<lv>" into a single dm name by doubling every literal "-" within the
VG/LV name components (e.g. vg "vdo-<uuid-with-dashes>" becomes
"vdo--<uuid-with-double-dashes>" in dmsetup ls output) -- the prefix match
was comparing against the unescaped VG name, which never matches. Escapes
the VG name the same way device-mapper does before matching.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
vdo.go's device-scoping and content-based identity logic (devicesArgs/vgExists/vgHasLV/pvVGName/runLVMCommand/removeOrphanedDMNodes) was never VDO-specific -- it's a general fix for LVM's default behavior breaking against a simplyblock HA volume's two redundant local device nodes. Extracted to github.com/simplyblock/atlas/lvm (PR #457) so a future pNFS striped-volume implementation (PR #456's design already plans the same pvcreate/vgcreate/lvcreate assembly pattern) doesn't rediscover the same bugs. This adopts the extracted package here and deletes the private copy; behavior is unchanged, confirmed by the existing pkg/spdk and pkg/util tests passing unmodified. The atlas-lib/lvm source lands here too, ahead of PR #457 merging -- both branches carry the identical commit content until #457 merges and this branch rebases past it, at which point the duplication collapses on its own. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ViL41RrgYVvQaTVShyDxsq
boddumanohar
added a commit
that referenced
this pull request
Aug 25, 2026
Extracts the general, non-VDO-specific LVM plumbing out of csi-driver/pkg/util/vdo.go (issue #277, PR #402) into a shared primitive: scoping every LVM command to an explicit device list via --devices, and answering PV/VG/LV identity questions from a device's own on-disk content rather than a name lookup. The seam exists because a simplyblock NVMe-oF HA volume surfaces as more than one local device node on the client, each presenting byte-identical backend content. LVM's default behavior -- a system-wide device scan, and existence checks keyed by name rather than by device -- cannot tell those device nodes apart. VDO hit this live as a genuine "duplicate PV" ambiguity, and separately as a name-based `vgs <name>` check reporting a VG as present when it had never actually been created on the device being asked about. There is no second real consumer's code to adopt yet: PR #456 (pnfs-design) is design-only today, but its striped-volume design explicitly plans the same pvcreate/vgcreate/lvcreate assembly pattern, including the identical partial-state recovery problem VDO already found and fixed (its own test plan's SI-03: "Re-running after a failure between vgcreate and lvcreate completes without duplicating"). This extracts the primitive now so vdo.go can become the first real consumer (adoption is a follow-up commit on PR #402), and future pNFS implementation work builds on it from day one instead of rediscovering the same bugs. Package surface: Inspector.VolumeGroup (content-based PV -> VG identity), Inspector.HasLogicalVolume (distinguishes a fully assembled stack from an orphaned partial create), Inspector.Rescan (device-scoped pvscan --cache), Inspector.Run (arbitrary LVM/dm command, --devices-scoped), DeviceScope (the --devices argument, generalized to a comma-joined multi-device list for a striped VG), and EscapeDMName/RemoveOrphanedDMNodes (device-mapper's dash-escaping and orphaned-node cleanup). All behind a Runner function type so callers can test the identity logic without lvm2 or a kernel present. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ViL41RrgYVvQaTVShyDxsq
boddumanohar
added a commit
that referenced
this pull request
Aug 26, 2026
Addresses PR review feedback: - Inspector was a bad name once the type grew a Run method that executes arbitrary mutating commands, not just answers identity questions. Renamed to Manager, matching this repo's own nvmeof.Connector precedent (name the type after what it does). - The package was too raw: every operation beyond the three identity checks (VolumeGroup, HasLogicalVolume, Rescan) required a caller to hand-build Run() argument lists -- exactly the duplication this package exists to prevent, and exactly what vdo.go had to do for pvcreate/vgcreate/lvcreate/vgchange/vgremove/lvextend/vgimportclone/ lvrename. Added named, purpose-built methods for all of them (CreatePhysicalVolume, CreateVolumeGroup, CreateVDOLogicalVolume, Activate/DeactivateVolumeGroup, RemoveVolumeGroup, ImportClonedVolumeGroup, RenameLogicalVolume, ExtendPhysicalVolume, ExtendLogicalVolumeByFreeSpace, LogicalVolumeSize, ExtendLogicalVolumeToSize, SetVDOFeatures), following the shape openebs/lvm-localpv's pkg/lvm uses (named operations, not a raw command-line builder) minus the Kubernetes-CRD-shaped parts of that package, which don't belong in atlas-lib per this repo's own convention. Run stays as the escape hatch for anything not covered. - Split the single lvm.go into one file per concern (lvm.go, identity.go, volume.go, vdo.go, clone.go, grow.go, dm.go), matching how nvmeof and nvme are already organized in this library, and matching the openebs package's own per-concern file layout. vdo.go's adoption of this expanded surface is a follow-up commit on PR #402. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ViL41RrgYVvQaTVShyDxsq
Follow-up to PR #457's rewrite of atlas-lib/lvm (Inspector -> Manager, plus named methods for every LVM operation VDO needs). Replaces every lvmInspector.Run(...) argument-building call site in vdo.go with the matching named method: CreatePhysicalVolume, CreateVolumeGroup, CreateVDOLogicalVolume, ActivateVolumeGroup, DeactivateVolumeGroup, RemoveVolumeGroup, ImportClonedVolumeGroup, RenameLogicalVolume, ListLogicalVolumes, ExtendPhysicalVolume, ExtendLogicalVolumeByFreeSpace, LogicalVolumeSize, ExtendLogicalVolumeToSize, SetVDOFeatures. Drops vdo.go's now-redundant local yn() helper (folded into the named CreateVDOLogicalVolume/SetVDOFeatures methods) and the raw lvs-based size-parsing/LV-listing logic now handled by LogicalVolumeSize/ ListLogicalVolumes. Behavior is unchanged -- same commands, same scoping, same fallback paths (DeactivateVDO's "not found" string match, RemoveVDO/DeactivateVDO's dmsetup fallback). Copies atlas-lib/lvm/*.go from the extract-lvm-atlas-lib branch (PR #457) into this worktree's copy of atlas-lib, since the two branches aren't rebased onto each other yet. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ViL41RrgYVvQaTVShyDxsq
boddumanohar
added a commit
that referenced
this pull request
Aug 26, 2026
Extracts the general, non-VDO-specific LVM plumbing out of csi-driver/pkg/util/vdo.go (issue #277, PR #402) into a new shared atlas-lib/lvm package, ahead of PR #456's pNFS design needing the same class of LVM assembly (pvcreate/vgcreate/lvcreate --stripes) and the same bugs VDO already found and fixed live: LVM's duplicate-PV ambiguity between a volume's redundant HA device nodes, name-based `vgs` lookups that lie about VG existence, and orphaned-VG detection after a partial pvcreate+vgcreate-but-not-lvcreate. lvm.Manager runs every LVM/dm-vdo command scoped to a fixed set of devices, and answers device-content identity questions about them, all behind a Runner function type (mirrors nvmeof.CommandRunner) so callers can test the identity logic without lvm2 or a kernel present. One file per concern, matching how nvmeof/nvme are already organized in this library: - identity.go: VolumeGroup (content-based PV -> VG identity lookup, not `vgs <name>`), ListLogicalVolumes, HasLogicalVolume (distinguishes a fully assembled stack from one orphaned by an interrupted create), Rescan (device-scoped `pvscan --cache`). - volume.go: CreatePhysicalVolume, CreateVolumeGroup, ActivateVolumeGroup, DeactivateVolumeGroup, RemoveVolumeGroup. - vdo.go: CreateVDOLogicalVolume, SetVDOFeatures. - clone.go: ImportClonedVolumeGroup, RenameLogicalVolume (resolving a byte-level clone/snapshot restore's PV/VG UUID collision). - grow.go: ExtendPhysicalVolume, ExtendLogicalVolumeByFreeSpace, LogicalVolumeSize, ExtendLogicalVolumeToSize. - dm.go: EscapeDMName / RemoveOrphanedDMNodes (device-mapper's dash-escaping and orphaned-node cleanup, generalized from VDO's own fix for the same class of failure). - lvm.go: Run, the escape hatch for anything not covered by a named method above, plus DeviceScope (the --devices argument, generalized to a comma-joined multi-device list -- VDO needs one device, a striped VG needs n). vdo.go's adoption of this package is a follow-up commit on PR #402. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ViL41RrgYVvQaTVShyDxsq
boddumanohar
added a commit
that referenced
this pull request
Aug 26, 2026
- Unexport DeviceScope -> deviceScope and EscapeDMName -> escapeDMName: neither has a caller outside this package (confirmed via grep across operator and csi-driver), so both were exposing an argument-building helper as public API rather than the named operation it exists to support. - Add ExtendVolumeGroup (vgextend), the counterpart to CreateVolumeGroup for growing an existing VG's device membership -- flagged as a gap in grow.go, which otherwise only extends a PV or an LV, not the VG itself. Needed by a striped VG that grows by adding a member, not just by resizing one already in place. ExtendLogicalVolumeToSize is used by csi-driver/pkg/util/vdo.go's GrowVDO on PR #402 (growing the VDO logical volume to match its pool's new size after ExtendLogicalVolumeByFreeSpace) -- not visible from this PR's own diff since that adoption lives on the sibling branch. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ViL41RrgYVvQaTVShyDxsq
Syncs vdo.go with atlas-lib/lvm's ExtendPhysicalVolume -> ExpandPhysicalVolume / ExtendLogicalVolumeByFreeSpace -> ExpandLogicalVolume rename (PR #457, commit 5ebffe9). No behavior change. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ViL41RrgYVvQaTVShyDxsq
Follows PR #457 commit 960cb06. No effect on this branch's own code -- vdo.go never referenced lvm.Runner or lvm.NewManagerWithRunner directly. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ViL41RrgYVvQaTVShyDxsq
noctarius
pushed a commit
that referenced
this pull request
Aug 28, 2026
Extracts the general, non-VDO-specific LVM plumbing out of csi-driver/pkg/util/vdo.go (issue #277, PR #402) into a new shared atlas-lib/lvm package, ahead of PR #456's pNFS design needing the same class of LVM assembly (pvcreate/vgcreate/lvcreate --stripes) and the same bugs VDO already found and fixed live: LVM's duplicate-PV ambiguity between a volume's redundant HA device nodes, name-based `vgs` lookups that lie about VG existence, and orphaned-VG detection after a partial pvcreate+vgcreate-but-not-lvcreate. lvm.Manager runs every LVM/dm-vdo command scoped to a fixed set of devices, and answers device-content identity questions about them, all behind a Runner function type (mirrors nvmeof.CommandRunner) so callers can test the identity logic without lvm2 or a kernel present. One file per concern, matching how nvmeof/nvme are already organized in this library: - identity.go: VolumeGroup (content-based PV -> VG identity lookup, not `vgs <name>`), ListLogicalVolumes, HasLogicalVolume (distinguishes a fully assembled stack from one orphaned by an interrupted create), Rescan (device-scoped `pvscan --cache`). - volume.go: CreatePhysicalVolume, CreateVolumeGroup, ActivateVolumeGroup, DeactivateVolumeGroup, RemoveVolumeGroup. - vdo.go: CreateVDOLogicalVolume, SetVDOFeatures. - clone.go: ImportClonedVolumeGroup, RenameLogicalVolume (resolving a byte-level clone/snapshot restore's PV/VG UUID collision). - grow.go: ExtendPhysicalVolume, ExtendLogicalVolumeByFreeSpace, LogicalVolumeSize, ExtendLogicalVolumeToSize. - dm.go: EscapeDMName / RemoveOrphanedDMNodes (device-mapper's dash-escaping and orphaned-node cleanup, generalized from VDO's own fix for the same class of failure). - lvm.go: Run, the escape hatch for anything not covered by a named method above, plus DeviceScope (the --devices argument, generalized to a comma-joined multi-device list -- VDO needs one device, a striped VG needs n). vdo.go's adoption of this package is a follow-up commit on PR #402. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ViL41RrgYVvQaTVShyDxsq
noctarius
pushed a commit
that referenced
this pull request
Aug 28, 2026
- Unexport DeviceScope -> deviceScope and EscapeDMName -> escapeDMName: neither has a caller outside this package (confirmed via grep across operator and csi-driver), so both were exposing an argument-building helper as public API rather than the named operation it exists to support. - Add ExtendVolumeGroup (vgextend), the counterpart to CreateVolumeGroup for growing an existing VG's device membership -- flagged as a gap in grow.go, which otherwise only extends a PV or an LV, not the VG itself. Needed by a striped VG that grows by adding a member, not just by resizing one already in place. ExtendLogicalVolumeToSize is used by csi-driver/pkg/util/vdo.go's GrowVDO on PR #402 (growing the VDO logical volume to match its pool's new size after ExtendLogicalVolumeByFreeSpace) -- not visible from this PR's own diff since that adoption lives on the sibling branch. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ViL41RrgYVvQaTVShyDxsq
boddumanohar
added a commit
that referenced
this pull request
Aug 28, 2026
Moves CreateOrAttachVDO, ResolveClonedVDO, DeactivateVDO, RemoveVDO, and GrowVDO out of csi-driver/pkg/util/vdo.go into this package, as CreateOrAttach, ResolveClone, Deactivate, Remove, and Grow. None of these functions reference a Kubernetes or CSI type: they orchestrate several lvm.Manager calls with judgment calls that are LVM/VDO-domain (when to reactivate vs. recreate, when to fall back to RemoveOrphanedDMNodes) rather than CSI-domain, so per this repo's own atlas-lib placement rule -- a node-level primitive belongs here, Kubernetes-shaped logic belongs in the consumer -- this is a better fit than where it sat. Also adds SetFeatures, a lvolID-keyed convenience wrapper over UpdateVolume, and DevicePath, so every exported function in this file is addressed by lvolID alone; the volume group/pool naming convention (volumeGroupPrefix, poolName) stays internal to this package rather than leaking to a caller. ResolveClone is now a thin wrapper over ResolveClonedVolumeGroup (already in lvm/clone.go) rather than its own copy of the rescan/probe/import/rename sequence. Two things changed in the move, both flagged in code review before this landed: - Logging: the original used k8s.io/klog directly. atlas-lib has no Kubernetes dependency anywhere, so this package exports Logger (a package-level *slog.Logger, nil-safe, matching the existing errs/deferrers.Logger precedent) instead. - Timeouts: the original's per-call context.WithTimeout budgets (120s/300s) were csi-driver-owned constants. This package now owns them as its own defaults; a caller can still tighten via its own ctx deadline. vdo.go's adoption of this package (deleting its own copies) is a follow-up commit on PR #402. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ViL41RrgYVvQaTVShyDxsq
Syncs atlas-lib/lvm with extract-lvm-atlas-lib (PR #457) through 74c745d: CreateLogicalVolume dispatch/Handles fixes, the physicalVolume->poolName rename, Extend->Expand renaming, Runner unexport, and -- the actual point of this commit -- CreateOrAttachVDO, ResolveClonedVDO, DeactivateVDO, RemoveVDO, and GrowVDO moving from this file into atlas-lib/lvm/vdo as CreateOrAttach, ResolveClone, Deactivate, Remove, and Grow. vdo.go is now thin wiring: one shared lvm.Manager, and one function per NodeStageVolume/NodeUnstageVolume/NodeExpandVolume concern, each delegating straight into the vdo package. The klog.Warningf calls and the vgName/poolLVName naming convention moved with the functions that used them; nothing here duplicates them anymore. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ViL41RrgYVvQaTVShyDxsq
boddumanohar
added a commit
that referenced
this pull request
Aug 28, 2026
Extracts the general, non-VDO-specific LVM plumbing out of csi-driver/pkg/util/vdo.go (issue #277, PR #402) into a new shared atlas-lib/lvm package, ahead of PR #456's pNFS design needing the same class of LVM assembly (pvcreate/vgcreate/lvcreate --stripes) and the same bugs VDO already found and fixed live: LVM's duplicate-PV ambiguity between a volume's redundant HA device nodes, name-based `vgs` lookups that lie about VG existence, and orphaned-VG detection after a partial pvcreate+vgcreate-but-not-lvcreate. lvm.Manager runs every LVM/dm-vdo command scoped to a fixed set of devices, and answers device-content identity questions about them, all behind a Runner function type (mirrors nvmeof.CommandRunner) so callers can test the identity logic without lvm2 or a kernel present. One file per concern, matching how nvmeof/nvme are already organized in this library: - identity.go: VolumeGroup (content-based PV -> VG identity lookup, not `vgs <name>`), ListLogicalVolumes, HasLogicalVolume (distinguishes a fully assembled stack from one orphaned by an interrupted create), Rescan (device-scoped `pvscan --cache`). - volume.go: CreatePhysicalVolume, CreateVolumeGroup, ActivateVolumeGroup, DeactivateVolumeGroup, RemoveVolumeGroup. - vdo.go: CreateVDOLogicalVolume, SetVDOFeatures. - clone.go: ImportClonedVolumeGroup, RenameLogicalVolume (resolving a byte-level clone/snapshot restore's PV/VG UUID collision). - grow.go: ExtendPhysicalVolume, ExtendLogicalVolumeByFreeSpace, LogicalVolumeSize, ExtendLogicalVolumeToSize. - dm.go: EscapeDMName / RemoveOrphanedDMNodes (device-mapper's dash-escaping and orphaned-node cleanup, generalized from VDO's own fix for the same class of failure). - lvm.go: Run, the escape hatch for anything not covered by a named method above, plus DeviceScope (the --devices argument, generalized to a comma-joined multi-device list -- VDO needs one device, a striped VG needs n). vdo.go's adoption of this package is a follow-up commit on PR #402. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ViL41RrgYVvQaTVShyDxsq
boddumanohar
added a commit
that referenced
this pull request
Aug 28, 2026
- Unexport DeviceScope -> deviceScope and EscapeDMName -> escapeDMName: neither has a caller outside this package (confirmed via grep across operator and csi-driver), so both were exposing an argument-building helper as public API rather than the named operation it exists to support. - Add ExtendVolumeGroup (vgextend), the counterpart to CreateVolumeGroup for growing an existing VG's device membership -- flagged as a gap in grow.go, which otherwise only extends a PV or an LV, not the VG itself. Needed by a striped VG that grows by adding a member, not just by resizing one already in place. ExtendLogicalVolumeToSize is used by csi-driver/pkg/util/vdo.go's GrowVDO on PR #402 (growing the VDO logical volume to match its pool's new size after ExtendLogicalVolumeByFreeSpace) -- not visible from this PR's own diff since that adoption lives on the sibling branch. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ViL41RrgYVvQaTVShyDxsq
boddumanohar
added a commit
that referenced
this pull request
Aug 28, 2026
Moves CreateOrAttachVDO, ResolveClonedVDO, DeactivateVDO, RemoveVDO, and GrowVDO out of csi-driver/pkg/util/vdo.go into this package, as CreateOrAttach, ResolveClone, Deactivate, Remove, and Grow. None of these functions reference a Kubernetes or CSI type: they orchestrate several lvm.Manager calls with judgment calls that are LVM/VDO-domain (when to reactivate vs. recreate, when to fall back to RemoveOrphanedDMNodes) rather than CSI-domain, so per this repo's own atlas-lib placement rule -- a node-level primitive belongs here, Kubernetes-shaped logic belongs in the consumer -- this is a better fit than where it sat. Also adds SetFeatures, a lvolID-keyed convenience wrapper over UpdateVolume, and DevicePath, so every exported function in this file is addressed by lvolID alone; the volume group/pool naming convention (volumeGroupPrefix, poolName) stays internal to this package rather than leaking to a caller. ResolveClone is now a thin wrapper over ResolveClonedVolumeGroup (already in lvm/clone.go) rather than its own copy of the rescan/probe/import/rename sequence. Two things changed in the move, both flagged in code review before this landed: - Logging: the original used k8s.io/klog directly. atlas-lib has no Kubernetes dependency anywhere, so this package exports Logger (a package-level *slog.Logger, nil-safe, matching the existing errs/deferrers.Logger precedent) instead. - Timeouts: the original's per-call context.WithTimeout budgets (120s/300s) were csi-driver-owned constants. This package now owns them as its own defaults; a caller can still tighten via its own ctx deadline. vdo.go's adoption of this package (deleting its own copies) is a follow-up commit on PR #402. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ViL41RrgYVvQaTVShyDxsq
boddumanohar
added a commit
that referenced
this pull request
Aug 28, 2026
Per review (the CHANGES_REQUESTED blocking this PR): every method that
named a device, a volume group, or a logical volume by bare string now
takes/returns one of three new value types instead:
- PhysicalVolume{DevicePath}
- VolumeGroup{Name}
- LogicalVolume{VolumeGroup, Name}
Passing a device path where a VG name belongs, or a VG where an LV is
expected, is now a compile error instead of an LVM failure discovered
at runtime. LogicalVolume carries its VolumeGroup rather than a bare
name so a caller pairing the wrong VG with an LV by hand is also a
type error, not a "volume group not found" surprise.
None of the three references a Manager: they are plain, comparable
values (== answers "same identity"), not handles, so the same value
works with any Manager instance -- keeping Manager itself stateless,
per the earlier stateless-contract discussion on this PR.
lvm/vdo's exported functions (CreateOrAttach, Deactivate, Remove, Grow,
ResolveClone, SetFeatures, DevicePath) keep their existing string-keyed
signatures unchanged -- VDO's own volume group/pool naming convention
already encapsulates the identity, so the typed values stay internal
to lvm/vdo/stack.go. csi-driver/pkg/util/vdo.go (PR #402) needs no
changes as a result: confirmed by rebuilding it against this commit's
atlas-lib/lvm with zero source edits.
Every call site across the package and its ~150 tests updated to
match; build/vet/test/lint/house-style gate all clean.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViL41RrgYVvQaTVShyDxsq
…olume Follows PR #457 commit f562d50. No changes needed to csi-driver/pkg/util/vdo.go -- vdo.CreateOrAttach/Deactivate/Remove/ Grow/ResolveClone/SetFeatures keep their existing string-keyed signatures; the new typed values stay internal to atlas-lib/lvm/vdo/stack.go. Confirmed by rebuilding with zero source edits on this branch. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ViL41RrgYVvQaTVShyDxsq
noctarius
added a commit
that referenced
this pull request
Aug 28, 2026
* feat(atlas-lib): add lvm package for device-scoped LVM commands Extracts the general, non-VDO-specific LVM plumbing out of csi-driver/pkg/util/vdo.go (issue #277, PR #402) into a new shared atlas-lib/lvm package, ahead of PR #456's pNFS design needing the same class of LVM assembly (pvcreate/vgcreate/lvcreate --stripes) and the same bugs VDO already found and fixed live: LVM's duplicate-PV ambiguity between a volume's redundant HA device nodes, name-based `vgs` lookups that lie about VG existence, and orphaned-VG detection after a partial pvcreate+vgcreate-but-not-lvcreate. lvm.Manager runs every LVM/dm-vdo command scoped to a fixed set of devices, and answers device-content identity questions about them, all behind a Runner function type (mirrors nvmeof.CommandRunner) so callers can test the identity logic without lvm2 or a kernel present. One file per concern, matching how nvmeof/nvme are already organized in this library: - identity.go: VolumeGroup (content-based PV -> VG identity lookup, not `vgs <name>`), ListLogicalVolumes, HasLogicalVolume (distinguishes a fully assembled stack from one orphaned by an interrupted create), Rescan (device-scoped `pvscan --cache`). - volume.go: CreatePhysicalVolume, CreateVolumeGroup, ActivateVolumeGroup, DeactivateVolumeGroup, RemoveVolumeGroup. - vdo.go: CreateVDOLogicalVolume, SetVDOFeatures. - clone.go: ImportClonedVolumeGroup, RenameLogicalVolume (resolving a byte-level clone/snapshot restore's PV/VG UUID collision). - grow.go: ExtendPhysicalVolume, ExtendLogicalVolumeByFreeSpace, LogicalVolumeSize, ExtendLogicalVolumeToSize. - dm.go: EscapeDMName / RemoveOrphanedDMNodes (device-mapper's dash-escaping and orphaned-node cleanup, generalized from VDO's own fix for the same class of failure). - lvm.go: Run, the escape hatch for anything not covered by a named method above, plus DeviceScope (the --devices argument, generalized to a comma-joined multi-device list -- VDO needs one device, a striped VG needs n). vdo.go's adoption of this package is a follow-up commit on PR #402. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ViL41RrgYVvQaTVShyDxsq * fix(atlas-lib/lvm): stop swallowing real command failures as not-found VolumeGroup and HasLogicalVolume folded every pvs/lvs failure -- a transient lock-contention or I/O error included -- into the same "nothing found" result a genuinely blank device or empty VG produces. The real adopting caller (csi-driver/pkg/util/vdo.go's CreateOrAttachVDO) checks `err != nil` right after each call expecting to catch exactly that class of failure, but the check was dead code: a transient lvs error read as "orphaned VG, zero LVs" and would drive CreateOrAttachVDO into vgremove -f against a real, valid volume. VolumeGroup now only treats pvs's own "no PV signature" text as a blank device (matching the text-based signal isAlreadyConnected reads for nvme-cli elsewhere in atlas-lib) and propagates everything else. HasLogicalVolume runs after VolumeGroup has already confirmed the VG exists on this device, so an lvs failure there is never a legitimate "empty VG" signal -- it now propagates unconditionally. Also wraps DeactivateVolumeGroup's and RemoveVolumeGroup's errors with operation/VG context, matching every sibling method in the file. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(atlas-lib/lvm): address review feedback on scope and coverage - Unexport DeviceScope -> deviceScope and EscapeDMName -> escapeDMName: neither has a caller outside this package (confirmed via grep across operator and csi-driver), so both were exposing an argument-building helper as public API rather than the named operation it exists to support. - Add ExtendVolumeGroup (vgextend), the counterpart to CreateVolumeGroup for growing an existing VG's device membership -- flagged as a gap in grow.go, which otherwise only extends a PV or an LV, not the VG itself. Needed by a striped VG that grows by adding a member, not just by resizing one already in place. ExtendLogicalVolumeToSize is used by csi-driver/pkg/util/vdo.go's GrowVDO on PR #402 (growing the VDO logical volume to match its pool's new size after ExtendLogicalVolumeByFreeSpace) -- not visible from this PR's own diff since that adoption lives on the sibling branch. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ViL41RrgYVvQaTVShyDxsq * refactor(atlas-lib/lvm): Extend -> Expand for the fixed-target grow ops Per review: ExtendPhysicalVolume and ExtendLogicalVolumeByFreeSpace always grow to a fixed target (the device's current full size, all newly available free space) with no partial-amount use case, so "Expand" fits better than "Extend." ExtendLogicalVolumeToSize and ExtendVolumeGroup keep "Extend" -- both take an explicit target (a byte size, a set of device paths) rather than growing to a fixed endpoint. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ViL41RrgYVvQaTVShyDxsq * refactor(atlas-lib/lvm): unexport Runner -> runner Per review: nothing outside this package injects a runner today, so the type name itself doesn't need to be exported. NewManagerWithRunner stays exported and still usable from outside the package regardless -- Go allows passing a matching func literal to a named func-type parameter whether or not the type name is exported, so the seam for a future consumer's tests is unchanged. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ViL41RrgYVvQaTVShyDxsq * Extracted vdo-specific code into its own sub-package * Updated the parameter set * quality gate run against atlas-lib/lvm * fix(atlas-lib/lvm): fix CreateLogicalVolume dispatch and Volume construction Three rough edges in the CreateLogicalVolume/LogicalVolumeDefinition provisioning-handler design (introduced in 8908e47/e642b9af): - CreateLogicalVolume hardcoded a lookup of volumeProvisioning["vdo"] rather than dispatching by asking each registered handler whether it Handles(def), so a second handler (striping, say) would never be reachable regardless of what it claimed to handle. Added Handles to the VolumeProvisioning interface (it existed only on the concrete vdo.volumeHandler, unreachable through the interface) and iterate registered handlers instead of hardcoding a key. Proven red first: a fake handler registered under a different name was silently ignored by the old hardcoded lookup. - vdo.volumeHandler.Handles required both Compression and Deduplication (&&), while CreateVolumeArgs contributed flags for either one (||) -- harmless only because Handles was never actually called before this fix. Now that CreateLogicalVolume dispatches through it, the mismatch would have meant a compression-only or deduplication-only volume silently got no VDO flags at all. Fixed Handles to match CreateVolumeArgs's condition, test-first. - vdo.Volume had every field unexported and no constructor, so nothing outside the vdo package (csi-driver, the actual intended caller of UpdateVolume) could construct one. Added vdo.NewVolume. Also renamed CreateLogicalVolume's physicalVolume parameter to poolName: it was never a physical volume, VDO's own call site already passed its pool LV name ("vdopool") through it, and lvcreate's <vg>/<pool> form only ever names a pool a handler creates alongside the logical volume, never a device. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ViL41RrgYVvQaTVShyDxsq * feat(atlas-lib/lvm/vdo): move the VDO stack lifecycle from csi-driver Moves CreateOrAttachVDO, ResolveClonedVDO, DeactivateVDO, RemoveVDO, and GrowVDO out of csi-driver/pkg/util/vdo.go into this package, as CreateOrAttach, ResolveClone, Deactivate, Remove, and Grow. None of these functions reference a Kubernetes or CSI type: they orchestrate several lvm.Manager calls with judgment calls that are LVM/VDO-domain (when to reactivate vs. recreate, when to fall back to RemoveOrphanedDMNodes) rather than CSI-domain, so per this repo's own atlas-lib placement rule -- a node-level primitive belongs here, Kubernetes-shaped logic belongs in the consumer -- this is a better fit than where it sat. Also adds SetFeatures, a lvolID-keyed convenience wrapper over UpdateVolume, and DevicePath, so every exported function in this file is addressed by lvolID alone; the volume group/pool naming convention (volumeGroupPrefix, poolName) stays internal to this package rather than leaking to a caller. ResolveClone is now a thin wrapper over ResolveClonedVolumeGroup (already in lvm/clone.go) rather than its own copy of the rescan/probe/import/rename sequence. Two things changed in the move, both flagged in code review before this landed: - Logging: the original used k8s.io/klog directly. atlas-lib has no Kubernetes dependency anywhere, so this package exports Logger (a package-level *slog.Logger, nil-safe, matching the existing errs/deferrers.Logger precedent) instead. - Timeouts: the original's per-call context.WithTimeout budgets (120s/300s) were csi-driver-owned constants. This package now owns them as its own defaults; a caller can still tighten via its own ctx deadline. vdo.go's adoption of this package (deleting its own copies) is a follow-up commit on PR #402. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ViL41RrgYVvQaTVShyDxsq * feat(atlas-lib/lvm): typed PhysicalVolume, VolumeGroup, LogicalVolume Per review (the CHANGES_REQUESTED blocking this PR): every method that named a device, a volume group, or a logical volume by bare string now takes/returns one of three new value types instead: - PhysicalVolume{DevicePath} - VolumeGroup{Name} - LogicalVolume{VolumeGroup, Name} Passing a device path where a VG name belongs, or a VG where an LV is expected, is now a compile error instead of an LVM failure discovered at runtime. LogicalVolume carries its VolumeGroup rather than a bare name so a caller pairing the wrong VG with an LV by hand is also a type error, not a "volume group not found" surprise. None of the three references a Manager: they are plain, comparable values (== answers "same identity"), not handles, so the same value works with any Manager instance -- keeping Manager itself stateless, per the earlier stateless-contract discussion on this PR. lvm/vdo's exported functions (CreateOrAttach, Deactivate, Remove, Grow, ResolveClone, SetFeatures, DevicePath) keep their existing string-keyed signatures unchanged -- VDO's own volume group/pool naming convention already encapsulates the identity, so the typed values stay internal to lvm/vdo/stack.go. csi-driver/pkg/util/vdo.go (PR #402) needs no changes as a result: confirmed by rebuilding it against this commit's atlas-lib/lvm with zero source edits. Every call site across the package and its ~150 tests updated to match; build/vet/test/lint/house-style gate all clean. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ViL41RrgYVvQaTVShyDxsq --------- Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> Co-authored-by: Christoph Engelbert (noctarius) <me@noctarius.com>
boddumanohar
added a commit
that referenced
this pull request
Aug 28, 2026
…plementation The document under this filename described a generalized, pluggable node-side-stack framework that was never built. Replace it with a design doc and test plan describing the actual mechanism PR #402 implements: StoragePool-level VDO compression/deduplication built on atlas-lib/lvm and atlas-lib/lvm/vdo (PR #457), node capability gating, and CSI wiring. Rename both files from design-node-volume-stack.md / test-plan-node-volume-stack.md to design-client-side-vdo-compression.md / test-plan-client-side-vdo-compression.md to match.
boddumanohar
added a commit
that referenced
this pull request
Aug 31, 2026
Tested against the real implementation (PR #402), not just raw LVM commands: two real PVCs with clientCompression/clientDeduplication on the same node, distinct checksummed data, node rebooted. Both VDO instances reattached cleanly via fresh NodeStageVolume calls (kubelet's own bookkeeping resets on reboot too) -- kvdo module usage count exactly 2, both VDOOperatingMode normal, both checksums matched exactly. Caveat found and documented: the two NodeStageVolume LVM command sequences happened to complete sequentially rather than genuinely overlapping, so LVM's internal command locking under truly concurrent vgchange/pvscan calls remains unexercised -- narrowed the open item accordingly rather than closing it outright. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
boddumanohar
added a commit
that referenced
this pull request
Aug 31, 2026
…ation Tested both clone paths against PR #402's real ResolveClonedVDO: a direct PVC-to-PVC clone and a snapshot restore, both scheduled onto the same node as their still-live source (the specific co-location scenario this finding warns about). Both correctly resolved via vgimportclone + lvrename, mounted cleanly with data matching the source exactly, and coexisted with the source and each other with independent VG identities and no cross-contamination. Also corrected the "Detection" section: the implementation ended up simpler than originally planned -- detection is unconditional and purely device-identity-based, not gated on VolumeContentSource, so no separate content-source plumbing was needed. Found and fixed one bug along the way: the collision-detection log message was picking up pvs's stderr WARNING: lines merged into its output instead of just the VG name -- harmless in this run, but fixed to parse cleanly. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
boddumanohar
added a commit
that referenced
this pull request
Aug 31, 2026
Deliberately reproduced against PR #402's real implementation: forcibly disconnected a VDO volume's NVMe-oF subsystem at the host level while the node stayed up, then deleted the pod. This exposed two real bugs (both now fixed on that branch): DeactivateVDO had no fallback for an unreachable device, and once added, the fallback's device-name matching didn't account for device-mapper's dash-escaping and matched nothing. With both fixed, cleanup is now fully automatic -- confirmed by reproducing the whole sequence a second time. Also documented an unplanned but valuable side observation: for ~19s after disconnect, cached reads/writes silently appeared to succeed before the real I/O failure surfaced, at which point VDO correctly fenced itself into read-only mode and ext4 independently aborted its journal -- both layers protected data correctly with no wiring needed from this design. Narrowed the remaining open item: the node-NotReady-and-rejoin path is still unverified (this test kept the node itself healthy throughout, only the storage connection was severed). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
boddumanohar
added a commit
that referenced
this pull request
Sep 1, 2026
Tested against the real implementation (PR #402), not just raw LVM commands: two real PVCs with clientCompression/clientDeduplication on the same node, distinct checksummed data, node rebooted. Both VDO instances reattached cleanly via fresh NodeStageVolume calls (kubelet's own bookkeeping resets on reboot too) -- kvdo module usage count exactly 2, both VDOOperatingMode normal, both checksums matched exactly. Caveat found and documented: the two NodeStageVolume LVM command sequences happened to complete sequentially rather than genuinely overlapping, so LVM's internal command locking under truly concurrent vgchange/pvscan calls remains unexercised -- narrowed the open item accordingly rather than closing it outright. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
boddumanohar
added a commit
that referenced
this pull request
Sep 1, 2026
…ation Tested both clone paths against PR #402's real ResolveClonedVDO: a direct PVC-to-PVC clone and a snapshot restore, both scheduled onto the same node as their still-live source (the specific co-location scenario this finding warns about). Both correctly resolved via vgimportclone + lvrename, mounted cleanly with data matching the source exactly, and coexisted with the source and each other with independent VG identities and no cross-contamination. Also corrected the "Detection" section: the implementation ended up simpler than originally planned -- detection is unconditional and purely device-identity-based, not gated on VolumeContentSource, so no separate content-source plumbing was needed. Found and fixed one bug along the way: the collision-detection log message was picking up pvs's stderr WARNING: lines merged into its output instead of just the VG name -- harmless in this run, but fixed to parse cleanly. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
boddumanohar
added a commit
that referenced
this pull request
Sep 1, 2026
Deliberately reproduced against PR #402's real implementation: forcibly disconnected a VDO volume's NVMe-oF subsystem at the host level while the node stayed up, then deleted the pod. This exposed two real bugs (both now fixed on that branch): DeactivateVDO had no fallback for an unreachable device, and once added, the fallback's device-name matching didn't account for device-mapper's dash-escaping and matched nothing. With both fixed, cleanup is now fully automatic -- confirmed by reproducing the whole sequence a second time. Also documented an unplanned but valuable side observation: for ~19s after disconnect, cached reads/writes silently appeared to succeed before the real I/O failure surfaced, at which point VDO correctly fenced itself into read-only mode and ext4 independently aborted its journal -- both layers protected data correctly with no wiring needed from this design. Narrowed the remaining open item: the node-NotReady-and-rejoin path is still unverified (this test kept the node itself healthy throughout, only the storage connection was severed). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
noctarius
added a commit
that referenced
this pull request
Sep 4, 2026
* Volume stack design iteration * Use the American spelling "unparsable" The §13 row read "unparseable", which is the British pattern of keeping the silent e before -able. American English drops it, which is what the house style wordlist already encodes for every comparable pair: useable to usable, moveable to movable, sizeable to sizable. "parse" is not a soft-c or soft-g stem, so it takes the same path. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Potential fix for pull request finding Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com> * Rewrite the issue #277 design doc and test plan around the shipped implementation The document under this filename described a generalized, pluggable node-side-stack framework that was never built. Replace it with a design doc and test plan describing the actual mechanism PR #402 implements: StoragePool-level VDO compression/deduplication built on atlas-lib/lvm and atlas-lib/lvm/vdo (PR #457), node capability gating, and CSI wiring. Rename both files from design-node-volume-stack.md / test-plan-node-volume-stack.md to design-client-side-vdo-compression.md / test-plan-client-side-vdo-compression.md to match. --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com> Co-authored-by: Manohar Reddy <manohar@simplyblock.io>
noctarius
force-pushed
the
main
branch
2 times, most recently
from
September 9, 2026 10:21
60dceb7 to
fbaabe4
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Implements the design in
operator/docs/designs/design-issue-277-client-side-compression.md(PR #398) and its test plan inoperator/docs/tests/test-plan-issue-277-client-side-compression.md, closing issue #277.Pool.spec.storageClassParameters.clientCompression/clientDeduplicationfields, independently switchable, mapped through toclient_compression/client_deduplicationStorageClass parameters.upsertStorageClassnow composes a VDO-capable topology requirement (simplyblock.io/vdo-capable=true) with the existing DHCHAP topology gate when either client-side parameter is set.csi-driver/pkg/util/vdo.go:CreateOrAttachVDO,ResolveClonedVDO,DeactivateVDO,RemoveVDO(with admsetup-based fallback for the orphaned-stack case documented in the design's spike log),GrowVDO,SetVDOFeatures. Each volume gets its own PV/VG/vdo-pool/LV stack via LVM (lvcreate --type vdo).nodeserver.gowiring:NodeStageVolumecreates/attaches VDO betweeninitiator.Connectand mount;NodeUnstageVolumedeactivates it before disconnecting the raw device;restageVolumereattaches (never recreates) on reconnect;NodeExpandVolumegrows the VDO stack before the filesystem resize;xfsStripeOptionsis skipped once VDO is in play;buildAccessibleTopologysurfaces the new node label.kmod-kvdo+vdoviansenter(addedhostPID: true), writing a marker file;newNodeServerreads it in the background and patches the node'ssimplyblock.io/vdo-capablelabel via a new RBACpatch/updategrant onnodes.csi-driver/deploy/image/Dockerfilenow installslvm2+vdo(amd64 only) and disables udev sync for LVM commands, both required to actually run LVM/VDO tooling from inside this container.Verified end-to-end on the live test cluster
Created a real
Poolwith bothclientCompressionandclientDeduplicationenabled, provisioned a PVC/Pod through the resulting StorageClass, and confirmed:client_compression/client_deduplication = "True"and anallowedTopologiesrequiringvdo-capable=true.kmod-kvdobuild (the other 3 nodes correctly advertisevdo-capable=false)./dev/mapper/vdo-<lvol-uuid>-...), withVDOCompression/VDODeduplicationbothenabled(lvs).vdostatssaving percent.NodeStageVolumecalls (kvdomodule usage count exactly 2, bothVDOOperatingMode: normal, both checksums matched). Caveat: the two stage sequences happened to run sequentially rather than genuinely overlapping, so LVM's internal locking under truly concurrentvgchange/pvscancalls remains unexercised.ResolveClonedVDOcorrectly ranvgimportclone+lvrename, both mounted cleanly with data matching the source exactly, and all three volumes (source, clone, restore) coexisted with independent VG identities and no cross-contamination.nvme disconnect) while the node and pod both stayed up, simulating "storage side disconnects while the node stays up." Kernel logs showed buffered reads/writes silently appearing to succeed for ~19s before the real failure surfaced, at which point VDO correctly fenced itself into read-only mode (and ext4 independently aborted its journal) rather than corrupting anything — Kubernetes reported the pod as healthy the whole time regardless. Deleting the pod afterward is now confirmed to clean up fully automatically, with no orphaned state left behind and no manual intervention needed.fsType=xfs(every prior spike usedext4) —mkfs.xfsran with no stripe-alignment flags (confirming thexfsStripeOptionsskip fires correctly), mounted cleanly with the existingnouuidflag intact, compression/dedup stayed enabled, and reattach-on-recreate worked identically toext4. No bugs found.asyncwrite policy: forcedvdo_write_policy=asyncexplicitly, wrote one file withfsync()and one without, then genuinely crashed the node viasysrq(immediate reboot, zero filesystem sync — confirmed via a new boot timestamp, not a graceful reboot that would have proven nothing). Thefsync()'d file survived with an exact checksum match; the non-fsync()'d file was lost entirely — the correct POSIX outcome. Resolves the design doc's open safety question:asynccorrectly honors flush/FUA durability end-to-end through NVMe-oF to the simplyblock backend.This process found and fixed nine real, hands-on-only-discoverable bugs beyond the original design (each is its own commit on this branch):
buildAccessibleTopologyonly added thevdo-capabletopology key when the label was alreadytrue— since CSINode's topology key set is captured once at plugin registration (seconds after pod start) while the label gets patched asynchronously afterward, the key was essentially never present at registration time, permanently breaking the topology gate until pod restart.spdkcsicontainer image had nolvm2installed at all (pvcreate: not found).lvm2alone wasn't enough —lvcreate --type vdoshells out tovdoformat, which ships in the separatevdopackage.lvcreate --type vdofailed with "device not cleared" — no udev daemon runs inside this container, so device-mapper's default udev-sync handshake never completes. Fixed withDM_DISABLE_UDEV=1.NodeUnstageVolumewas calling the destructiveRemoveVDO(vgremove) on every routine unstage, not only when the volume was actually being deleted — an ordinary pod delete+recreate on the same node silently destroyed all VDO-backed data. Fixed by adding a non-destructiveDeactivateVDO(vgchange -anonly) for this path.GrowVDO'slvextend -l100%FREEused the absolute (not additive) percentage form, which computed a target smaller than the pool's current size on every real device resize. Fixed with-l+100%FREE.pvVGNametrustedpvs's combined stdout+stderr output wholesale, so a duplicate-PVWARNING:line from a real clone test polluted both the identity comparison and the resulting log message. Harmless in that run (the dirty string still differed from the target either way), but fixed to parse just the actual field.DeactivateVDO(the non-destructive replacement from bug updated the api create param to include cr objects #5) had no fallback at all for an unreachable backing device —vgchange -anfailed identically to howvgremovefailed in the original spike, on every one of 18 retries, until kubelet gave up and force-removed the pod anyway, leaving the orphaned dm-vdo stack permanently stuck. Fixed by adding the samedmsetup removefallbackRemoveVDOalready had.vdo-<uuid>→vdo--<uuid-with-double-dashes>indmsetup lsoutput), so it matched nothing and the stack stayed orphaned even with the fallback wired in. Fixed to escape the VG name the same way device-mapper does before matching.Deliberate scope decisions (see design doc for full rationale)
ResolveClonedVDO's clone-collision detection is driven purely by the device's actual on-disk VG identity, not by threadingVolumeContentSourcethrough the volume context as the design doc originally proposed — simpler and correct either way, so no new plumbing was added.GrowVDO's signature grows to the pool's new physical capacity (-l100%FREE/+100%FREE, then matching logical size) rather than taking an explicitnewSizeparameter, matching the100%FREEconvention already used at creation time.csi-driver/pkg/util/vdo.goper the design doc, notatlas-lib(raised as an idea mid-discussion, never confirmed).lvm2/vdoare installed in this branch-taggedDockerfile, not the sharedbase_image/Dockerfile_base— rebuilding that shared, cross-branch tag for an in-progress feature felt like the wrong blast radius; worth promoting there once this lands.vdohas no aarch64 build in the configured repos, so it's installed amd64-only; client-side VDO is x86_64-only for now, matching every host used in this design's validation.New finding, now fixed (separate, pre-existing gap — fixed here for the VDO case)
The CSI driver's
CreateVolumedid not populatePersistentVolume.spec.nodeAffinity, so once a PVC was bound, a pod using it could be rescheduled to any node — not just the one the topology gate originally selected. This was harmless before (a raw NVMe-oF connection works identically from any node) but is now materially important, since VDO state is node-local. Reproduced live during this branch's original verification: deleting and recreating the pod (not the PVC) let it land on a non-vdo-capablenode, where it correctly failed to mount.After rebasing onto
origin/main, this branch picked up an upstream fix for the same underlying issue (#403) for the DHCHAP topology case (dhchapNodeLabelParam/dhchapAllowedNodeSegment). This PR adds the mirroringvdoCapableSegmentfor the VDO case:CreateVolumenow merges avdo-capabletopology segment intoAccessibleTopologywheneverclient_compression/client_deduplicationis set, exactly the same way the DHCHAP segment is merged, soexternal-provisionercorrectly pinsPersistentVolume.spec.nodeAffinity. Covered by new unit tests (TestVDOCapableSegment); not yet re-verified live post-rebase (the original live reproduction predates this fix).Explicitly out of scope for this PR
SetVDOFeatures(live compression/dedup toggle) is implemented but not wired into any update path — v1 non-goal per the design doc.upsertStorageClasscreate-only bug (bug: upsertStorageClass is create-only — Pool StorageClassParameters edits on an existing Pool silently no-op #401) — orthogonal, tracked separately.NodeUnstageVolumeon the original node specifically after that node goes NotReady and later rejoins (as opposed to the connection just being severed while the node stays healthy, which is what was tested) remains unverified — a different code path.vgchange/pvscancalls actually racing at the LVM level remain unexercised — the multi-instance reboot test's two stage sequences happened to run sequentially rather than overlapping.All items originally flagged as open in the test plan (XFS-on-VDO, crash-consistency of
async, clone/snapshot resolution, multi-instance-across-a-reboot, stale-state-after-unclean-disconnect) are now verified — see above.ensureDeviceConnected(block-volume reconnect) is untouched — every VDO example in the design is filesystem-mode; block-mode + VDO was never designed or tested.NodeExpandVolume/GrowVDOisn't fully idempotent against a redundant re-invocation after the volume is already at its target size — logs a scary-looking but harmless error on kubelet's post-success reconciliation retry. Worth a follow-up polish pass.Rebased onto origin/main
Rebased onto current
origin/main(37 commits of drift), all 12 original commits preserved plus onefollow-up commit regenerating the CRD/
install.yamlmanifests. Notable upstream changes this branch nowsits on top of, and how they were reconciled:
Pool→StoragePoolrename (CRD, controller, files all renamed on main): git's rename detectionfollowed this cleanly;
PoolReconciler→StoragePoolReconciler,pool_types.go→storagepool_types.go,simplyblockpool_controller.go→simplyblockstoragepool_controller.go. Thenew
ClientCompression/ClientDeduplicationfields and their wiring are all still present and intactpost-rebase.
upsertStorageClass→createStorageClassIfNotExists, now backed by new+k8s:immutableCELmarkers on
StorageClassParameters/DHCHAP: this looks like it resolves the semantic core of bug: upsertStorageClass is create-only — Pool StorageClassParameters edits on an existing Pool silently no-op #401(the create-only/no-drift-detection bug) via a different, arguably better mechanism — preventing the
edit at the API/admission level instead of detecting drift after the fact. Not actioned further in this
PR; flagging for whoever triages bug: upsertStorageClass is create-only — Pool StorageClassParameters edits on an existing Pool silently no-op #401.
dhchapNodeLabelParam/dhchapAllowedNodeSegment: main already fixes the "PVnodeAffinityisn'tpopulated" gap flagged above under "New finding" — but only for the DHCHAP topology case, not the VDO
one. This PR's
vdo-capabletopology requirement was merged alongside it increateStorageClassIfNotExists(both compose into oneTopologySelectorTermso they're ANDed, notORed, when a pool needs both). A follow-up commit on this branch (
vdoCapableSegment) closes the samegap for VDO specifically, mirroring
dhchapAllowedNodeSegmentexactly — see "New finding, now fixed"above.
FilesystemStorageClassParameter (defaultxfs, from unrelated upstream commits) merged cleanlyinto
mergeStorageClassParameters; not yet re-verified end-to-end against the VDO+XFS wiring now thatxfsis the operator's own default rather than something the original live testing had to force byhand — the underlying logic is unchanged, but this specific path wasn't re-run live post-rebase.
go build ./...andgo test ./pkg/util/... ./pkg/spdk/...(csi-driver) are clean post-rebase.golangci-lintshows only pre-existinggoconstfindings in unrelated test files, none touching any filethis PR modifies.
Why draft
Unit tests from the test plan are not written yet. Draft status reflects that, not any doubt about the functional behavior above — that part is now hands-on verified against a real cluster, real NVMe-oF-backed lvols, and real compression/dedup savings.
Test plan
go build ./...clean for bothoperatorandcsi-drivermodulesgo test ./...—pkg/spdk,pkg/utilpass;internal/controller's envtest-based suite requires a localkubebuilderbinary not present in this environment (pre-existing, unrelated);e2epackage requires a live e2e cluster fixture (pre-existing, unrelated)vdo.goand the new wiring (test plan Sections 1-7)🤖 Generated with Claude Code