Date: 2026-03-14
Sources: GitHub issues and PRs from kmsvnc (25 issues), gpu-screen-recorder, kms-screenshot, RustDesk, FFmpeg
| Issue | Format | GPU | Status | Impact |
|---|---|---|---|---|
| #1 | AR30 (10-bit) | AMD 6900XT | ⛔ Open since May 2023 | KDE Wayland uses 30-bit by default |
| #22 | AB48 (16-bit HDR) | AMD RX 550 | ⛔ Open | Fedora 42, kernel 6.14 |
| #24 | AB30 (10-bit) | Nvidia GTX | ⛔ Open | Nvidia BLOCK_LINEAR tiling |
| #7 | XRGB misdetection | AMD | 🔧 Resolved | Blue tint, incorrect channel swap |
Lesson for libdrmtap: Modern compositors (KDE, GNOME with HDR) output 10-bit (AR30/XR30) and 16-bit (AB48) by default. Supporting only XRGB8888 is insufficient for 2024+. User tried KWIN_DRM_PREFER_COLOR_DEPTH=24 to force 8-bit but it didn't work.
"weston uses two framebuffer IDs and flips between them for each frame. kmsvnc will use the currently active framebuffer ID. You will see the screen updates every 2 key presses instead of every key press."
Root cause: kmsvnc captures the framebuffer that was active at startup. Weston does page-flipping (alternates between 2 FBs). When the compositor switches to the other FB, kmsvnc keeps reading the old one.
Lesson for libdrmtap: Implement drmModeGetPlane() on every frame (like FFmpeg kmsgrab) to detect fb_id changes. Never cache the framebuffer ID.
"moving the mouse to the far right of the VNC client window puts the cursor on the right side of the right-hand monitor (crtc 80) instead of the right side of the left-hand monitor (crtc 77)"
Root cause: With 2 monitors at 1920px, the CRTC reports coordinates in the combined space (0-3840), but the cursor plane only has 0-1920. The conversion doesn't adjust.
Lesson for libdrmtap: Cursor coordinate space is relative to CRTC, not plane. In multi-monitor setups, mapping must be done correctly.
"I am trying to get this to work on a vkms device. [...] DRM ioctl error -1 on line 669"
Root cause: vkms doesn't support DRM_IOCTL_GEM_FLINK (the ioctl kmsvnc uses for "dumb buffer" path). The dumb buffer path fails with vkms.
Solution: Use drmPrimeHandleToFD instead of GEM_FLINK for vkms. The issue also confirms that on ArchLinux+Nvidia, vkms shows no active planes.
Lesson for libdrmtap: For vkms testing, use Prime pipeline (not dumb). Also, vkms needs something rendering to it (compositor) to have active planes.
"handles 0 0 0 0" without root, "handles 1 0 0 0" with
setcap cap_sys_admin=ep
Technical details from log:
- Driver: amdgpu
- Modifier:
AMD:GFX10_RBPLUS,GFX9_64K_R_X,DCC,DCC_RETILE,DCC_INDEPENDENT_128B,DCC_MAX_COMPRESSED_BLOCK=128B,DCC_CONSTANT_ENCODE,PIPE_XOR_BITS=4,PACKERS=3 - Without CAP_SYS_ADMIN:
handles 0 0 0 0→ kernel hides the handles - With CAP_SYS_ADMIN:
handles 1 0 0 0→ enabled
Lesson for libdrmtap: Confirms our API must detect handles[0] == 0 and activate the helper automatically.
"I noticed that in one buffer display data are offset by 512 bytes"
An ARM developer (iMX8) submitted 3 patches:
- Multiple framebuffer support (double-buffering)
- Fix for starting data address (offset)
imx-dcssdriver as new backend
Lesson for libdrmtap: ARM SoCs have arbitrary offsets in framebuffers. Direct mapping mmap(fd, 0) doesn't always start at byte 0.
"va operation error 0x14 the requested function is not implemented"
Intel CherryView (Atom x5-Z8330) with the old i965 driver fails on vaGetImage. Old VAAPI doesn't support the operation.
Lesson for libdrmtap: Fallback needed when VAAPI reports "not implemented". Don't assume Intel always = functional VAAPI.
All planes have CRTC 0 FB 0. Monitor is connected but compositor isn't using that card/CRTC.
Lesson for libdrmtap: There can be multiple DRM devices. If all planes on a device are empty, try another /dev/dri/cardN.
"Support Raspberrypi VC4 Broadcom T tiled"
Mentions an existing project: drm-vc4-grabber (Rust) that takes screenshots on RPi by decoding Broadcom's T-tiling.
Lesson for libdrmtap: RPi needs its own deswizzle (T-tiled, different from Intel X-tiled or Nvidia).
"The captured image may require rotation according to panel orientation property of the DRM connector"
References kernel documentation on panel orientation property.
Lesson for libdrmtap: Tablets and some laptops have rotated panels (90°, 180°, 270°). The framebuffer is captured "sideways". Must read the connector's orientation property.
"going from tty <-> games frontend <-> game/emulator, the video is always lost on the <->. I suspect this happens because framebuffers are destroyed and new ones are created"
Root cause: DRM context switch. When a fullscreen OpenGL app creates new framebuffers, the previous FB_ID is destroyed. kmsvnc loses the reference.
Lesson for libdrmtap: Implement automatic reconnection when framebuffer changes. Detect new fb_id on each frame (like FFmpeg).
"
SYS_pidfd_getfdsyscall available since Linux 5.6" + "fourcc_mod_is_vendormacro added 11 months ago, even ubuntu22.04 can not support"
Ubuntu 20.04 (kernel 5.4) doesn't have pidfd_getfd or new libdrm macros.
Lesson for libdrmtap: Define minimum supported kernel. Provide #ifndef fallbacks for libdrm macros that don't exist in older versions.
Weston already uses neatvnc. Suggests more permissive license (MIT vs GPL).
Lesson for libdrmtap: MIT/BSD license for maximum adoption. GPL is a problem for RustDesk (Apache 2.0) and proprietary projects.
gsr_kms_client_initcreates socketpair (not named Unix socket)- They replace file-backed socket with socketpair for security
- DMA-BUF FDs are passed via SCM_RIGHTS in socket messages
- NixOS: Binary paths not resolvable for
setcapdue to Nix store - Flatpak: Helper can't be found from the sandbox
- Polkit: Repetitive popup every time recording starts
"Unable to map DMABUF exported memory to CPU visible buffer: Permission denied"
Nvidia on Wayland has additional restrictions for mmapping exported DMA-BUFs.
Affected projects: kmsvnc, kms-screenshot
Frequency: ~30% of all issues
Problematic formats: AR30, AB30, AB48, ABGR16161616
Cause: HDR/10-bit compositors by default in 2024+
Affected projects: kmsvnc, gpu-screen-recorder, FFmpeg
Frequency: ~20% of issues
Cause: kernel 5.x+ hides handles without CAP_SYS_ADMIN
Standard solution: setcap cap_sys_admin+ep on helper
Affected projects: all
Frequency: ~15% of issues
Variants: Intel X-TILED, Intel CCS, Nvidia BLOCK_LINEAR, AMD DCC, Broadcom T-TILED
No universal solution: each GPU needs its own code path
DRM_FORMAT_MOD_VENDOR_NVIDIA is 0x03 (drm_fourcc.h:470). gpu_nvidia.c tested
0x10, which is not a vendor at all: it is the low byte of the block-linear encoding.
So the Nvidia branch never matched a real Nvidia modifier, and every Nvidia-vendor
scanout fell through to the linear memcpy and was reported as a successfully
converted frame -- the same fail-open shape as #4 below, one vendor over.
Two things made it survive: the deswizzler was reached only on hardware nobody here
runs the CPU path on, and a unit test asserted the behaviour of 0x10, so the test
certified the wrong constant rather than catching it.
Upstream definitions, for anyone verifying the encoding rather than trusting this
page: drm_fourcc.h
carries DRM_FORMAT_MOD_VENDOR_NVIDIA and the fourcc_mod_code(vendor, val) macro
that packs the vendor into bits 56-63; the kernel copy is
include/uapi/drm/drm_fourcc.h.
Rule: a vendor byte is modifier >> 56 compared against a DRM_FORMAT_MOD_VENDOR_*
macro, never a literal. If a literal appears, check it against drm_fourcc.h before
trusting any test that exercises it. And a decoder that cannot decode must return
-ENOTSUP from the branch itself, not travel through an allocation first: that path
returned -ENOMEM under memory pressure and hid the real reason.
Affected projects: libdrmtap (our own!)
Frequency: 100% on Intel Gen12+ with helper mode
GPU: Intel UHD (TGL/ADL) with I915_FORMAT_MOD_Y_TILED_CCS
Cause: dumb_mmap of CCS framebuffers returns compressed data that
cannot be CPU-deswizzled. Only EGL import via DMA-BUF fd works.
Impact: Silent process crash (out-of-bounds read in deswizzle)
Root cause chain:
- Helper does
MODE_MAP_DUMB+mmap→ gets CCS-compressed pixels - Sends raw CCS data over socket with
modifier = CCS - Parent has
dma_buf_fd = -1(V2 protocol: no SCM_RIGHTS) → EGL skipped - CPU
deswizzle_intel_y_tiled()runs on CCS data → out-of-bounds → crash
Why it wasn't caught before:
- vkms (test VM): modifier = LINEAR → deswizzle skipped entirely → no crash
- Standalone test (root): has
dma_buf_fd >= 0→ EGL GPU deswizzle → correct image - Only triggers when: Intel CCS + helper (non-root) + V2 protocol
Fix applied: CCS modifiers (0x05-0x08) return -ENOTSUP from drmtap_deswizzle().
gpu_auto_process() handles this by setting modifier to LINEAR (raw pixels, garbled but no crash).
Proper fix (shipped): the V3 helper protocol now passes the DMA-BUF fd via SCM_RIGHTS for
non-linear modifiers, so the parent uses the EGL/GLES2 GPU detile path (gpu_egl.c) instead of
CPU deswizzle. V3 (zero-copy) is the default; the V2 pixel copy is the fallback when DMA-BUF
export is not possible.
Affected projects: kmsvnc, RustDesk
Frequency: ~10% of issues
Cause: cursor coordinate space vs CRTC vs plane
Affected projects: kmsvnc
Frequency: low but critical
Cause: caching fb_id instead of refreshing it each frame
Affected projects: kmsvnc, FFmpeg
Frequency: ~10%
Cause: new APIs (GetFB2, pidfd_getfd, format modifiers)
Solution: #ifdef guards + fallback paths
From reading all these issues, the library MUST:
- Support AR30/XR30 (10-bit) and XR48/AR48 (16-bit HDR) — done (#16). HDR10 (PQ) scanouts are decoded and tone-mapped to SDR: PQ EOTF, BT.2020 → BT.709, a peak-aware highlight curve and sRGB, for
AR30/XR30and 16-bitXR48/AR48/XB48/AB48, in the CPU and EGL paths, gated on the connectorHDR_OUTPUT_METADATA. Plain SDR 10-bit is reduced directly;P010(overlay YUV) and HLG fall back - Refresh
plane->fb_idon every frame (never cache) — done (drm_grab.cre-readsdrmModeGetPlane()each frame) - Detect
handles[0] == 0and activate helper automatically — done (drm_grab.cfalls back to the privileged helper onhandles[0] == 0/EACCES) - Handle non-zero offsets in mmap (ARM SoCs)
- Have fallback when VAAPI reports "not implemented"
- Try multiple
/dev/dri/cardNif the first has no active planes — done (drmtap.cscanscard0..card15and keeps the device driving the MOST active CRTCs, falling back to any device with KMS resources;drmtap_list_devices()reaches the others) - Read
panel orientationproperty from connector - Define minimum kernel (5.6+) and add
#ifndefguards - Use MIT/BSD license (not GPL) — done (libdrmtap is MIT)
- Handle reconnection when framebuffers are destroyed and recreated
- Implement Prime pipeline (not GEM_FLINK) for vkms compatibility — done (
drmPrimeHandleToFD; vkms is used for CI-friendly synthetic scanout) - Document that NixOS/Flatpak need special configuration for the helper
- Do NOT implement damage tracking — FB_DAMAGE_CLIPS is for writers (compositors), not readers. Integrators (RustDesk, VNC) do their own frame diff
- Support continuous capture as a simple polling loop — no callbacks, no epoll required
- Protect
drmtap_ctxwithpthread_mutex_tfor thread safety (or document one-ctx-per-thread pattern) — done - Verify coexistence: multiple DRM readers (NoMachine, Sunshine, kmsvnc) can capture simultaneously
- Helper binary path must be configurable (compile-time + runtime) — done (
ctx->helper_pathfirst, then/usr/libexec/drmtap-helperand/usr/local/libexec/...). Note:libdrmtap-sysinstead embeds and statically compiles the helper sources - Set
FD_CLOEXECon received DMA-BUF fds to prevent leaking to child processes — done - Handle helper crash recovery — auto-respawn if helper dies mid-capture — done
- Never CPU-deswizzle CCS modifiers —
dumb_mmapreturns compressed data, only EGL can deswizzle. Return-ENOTSUPand require DMA-BUF fd + EGL (commit 7c7ce27) - V3 protocol: pass DMA-BUF fd via SCM_RIGHTS for non-linear modifiers so parent can use EGL GPU deswizzle — shipped and now the default zero-copy capture path; Intel CCS detiles correctly in unprivileged/helper mode, with the V2 pixel copy as the fallback
- Treat the convert descriptor as untrusted IPC input —
drmtap_convert_dmabuf()receives geometry and an fd from the other side of a process boundary. A descriptor claiming a frame larger than the fd actually backs made the CPU fallback mmap and read past the buffer, faulting the unprivileged converter with SIGBUS (a DoS). The fix requires a genuine DMA-BUF (a non-dma-buf fd is rejected; an immutable dma-buf cannot be truncated mid-read) and bounds the read against the buffer size fromlseek(reliable across kernels, unlikefstatwhich reports 0 for a dma-buf before Linux 5.3), failing closed when the size is unknown. Discovered 2026-07-20 by thefuzz/fuzz_convert.clibFuzzer harness (a descriptor whose declared frame exceeds the fd size) and fixed in 0.4.12; the harness is committed and guards against a regression - Read the vendor byte against
drm_fourcc.h, never a literal —gpu_nvidia.ctested0x10whereDRM_FORMAT_MOD_VENDOR_NVIDIAis0x03, so the branch was dead and Nvidia modifiers fell through to a linear copy reported as success. A unit test asserted the wrong constant. Fixed in 0.5.3 - Close EVERY GEM handle
drmModeGetFB2mints, not justhandles[0]— it creates a fresh handle per plane on every call, and only plane 0 is ever used here. On a scanout whose planes live in separate BOs (CCS is wherenum_planes >= 2comes from) the others leaked one handle per grab. Deduplicate first: planes sharing a BO come back as the SAME handle and a second close would free one still owned. Fixed in 0.5.3