{"slug": "nvidia-drivers-on-debian-13-trixie", "title": "Nvidia drivers on Debian 13 Trixie?", "summary": "A Proxmox VE 9 user running kernel 7.0 documented a working procedure for passing NVIDIA Tesla A2 GPUs through to privileged LXC containers with CUDA support on Debian 13 Trixie, pinning the NVIDIA driver to version 615 after version 590 failed to build against that kernel branch. The guide instructs blacklisting Nouveau, installing proxmox-headers-7.0 to match the running kernel, enabling the contrib component in Debian's deb822 sources, adding NVIDIA's cuda-keyring repository, and installing the nvidia-open driver with the versioned cuda-toolkit-13-4 metapackage rather than Debian's nvidia-cuda-toolkit. The write-up credits Claude with substantial involvement in resolving the driver and kernel issues.", "body_md": "So, a couple months later and I actually found myself in somewhat of the same situation you were in. I’ve got several Nvidia Tesla A2s in my Proxmox servers, and I wanted my LXCs to have access to them with Nvidia drivers and CUDA support.\n\nI’m currently running PVE9, on Kernel 7.0.\n\nHere’s what worked for me, admittedly with quite a bit of Claude’s involvement as I ran into driver+kernel issues and didn’t want to get too into the weeds myself. Apologies, as I pulled it from my internal Wiki article I made a month ago so versions may have updated since writing.\n\n## \nNVIDIA A2 Passthrough to Privileged LXC\n\nApplies to all Proxmox nodes with an A2 installed.\n\n## \n\nReference: [NVIDIA Tesla driver installation guide (Debian)](https://docs.nvidia.com/datacenter/tesla/driver-installation-guide/debian.html#local-repository-enablement)\n\n1. \nBlacklist Nouveau so it doesn’t grab the card first: \n\n```\necho -e \"blacklist nouveau\\noptions nouveau modeset=0\" > /etc/modprobe.d/blacklist-nouveau.conf\nupdate-initramfs -u\nreboot\n```\n\n2. \nInstall headers matching the **currently running** kernel — Proxmox\n doesn’t ship a`linux-headers-generic` virtual package, so target the\nbranch your active kernel belongs to:\n \n\n```\napt install proxmox-headers-$(uname -r | cut -d. -f1,2)\n```\n\n For `7.0.14-11-pve` , this resolves to`proxmox-headers-7.0` . Confirm with\n `uname -r` first if you’re not sure which kernel is currently booted —\n with multiple PVE kernels installed across update cycles,`apt install linux-headers-generic` has no candidate, and installing headers for a\n kernel other than the running one won’t let DKMS build the module. (If\n `uname -r` ever reports a non-`-pve` suffixed kernel, e.g. a stock Debian\n `6.12.107-1` build, that branch doesn’t follow this naming — install\n `linux-headers-amd64` instead.)\n3. \nMake sure `contrib` is enabled. Proxmox doesn’t ship\n `software-properties-common` , so`add-apt-repository` isn’t available —\n and Debian 13’s default`/etc/apt/sources.list.d/debian.sources` (deb822\n format) usually already includes`contrib non-free-firmware` out of the\nbox. Check first:\n \n\n```\ncat /etc/apt/sources.list.d/debian.sources\n```\n\n If `Components:` doesn’t already include`contrib` , add it (repeat for\n every stanza in the file, e.g. the`trixie-security` one too):\n \n\n```\nsed -i 's/^Components:\\(.*\\)$/Components:\\1 contrib/' /etc/apt/sources.list.d/debian.sources\napt update\n```\n\n You don’t need Debian’s `non-free` component for this — the CUDA driver\n comes from NVIDIA’s own repo via`cuda-keyring` in the next step, not\n from Debian’s`non-free` . Verify`contrib` is indexed:\n \n\n```\napt-cache policy | grep -A1 contrib\n```\n\n4. \nAdd the proprietary NVIDIA repo: \n\n```\nwget https://developer.download.nvidia.com/compute/cuda/repos/debian13/x86_64/cuda-keyring_1.1-1_all.deb\ndpkg -i cuda-keyring_1.1-1_all.deb\napt update\n```\n\n5. \nPin the GPU-compatible driver version. **`590` failed to build against\nthis kernel branch — see the troubleshooting note below — use `615`\ninstead** (or whatever the current working pin turns out to be at the\n time you’re reading this; check with`apt list -a nvidia-driver-pinning-*`\nif in doubt):\n \n\n```\napt install nvidia-driver-pinning-615\n```\n\n6. \nInstall the driver and CUDA toolkit. Use the versioned `cuda-toolkit-*`\n metapackage, not`nvidia-cuda-toolkit` — that name belongs to a\n different, Debian-maintained package that lives in Debian’s`non-free`\n component, not NVIDIA’s own repo, and won’t resolve with only`contrib`\n enabled. Check`apt-cache search cuda-toolkit` for the current versioned\noptions and pin one explicitly, the same way you pinned the driver:\n \n\n```\napt -V install nvidia-open cuda-toolkit-13-4\n```\n\n Watch the DKMS build output at the end of this command — it builds the\n kernel module for every installed kernel with matching headers present,\n and a compile failure here is what the troubleshooting note below walks\nthrough.\n7. \n**Run this exact sequence, with the same pinned driver (615) and toolkit\n(13-4) versions, on all three nodes.** Since containers may end up\n scheduled on any node via the`gpu-a2` resource mapping, a driver\n mismatch between nodes will surface as a container-vs-host version error\n even though each host is internally fine. The version pins make it easy\n to keep this in sync going forward — just re-run`apt install nvidia-driver-pinning-<version>` and pick the matching`cuda-toolkit-*`\non each node when you upgrade, rather than tracking versions by hand.\n\n### `page_free`/` zone_device_page_init` errors\n\nIf `apt -V install nvidia-open ...` fails during the `nvidia-kernel-open-dkms`\n\nstep with compiler errors like:\n\n``` js\nerror: 'const struct dev_pagemap_ops' has no member named 'page_free'\nerror: too few arguments to function 'zone_device_page_init'\n```\n\nthis is a genuine kernel-API incompatibility, not a headers-mismatch problem\n\n— it happens the same way regardless of which kernel/headers pair DKMS\n\nbuilds against. It means the pinned driver branch predates a kernel-side\n\nAPI change (a `struct page` → `struct folio` conversion in the\n\n`dev_pagemap_ops` memory-management code) that this Proxmox kernel branch\n\nalready has. The fix is a newer driver branch with updated kernel-compat\n\ncode, not anything on the headers/kernel side:\n\n```\napt list -a nvidia-driver-pinning-*\n```\n\nPick the newest available branch, clean up the broken partial install, and\n\nretry:\n\n```\napt remove --purge nvidia-driver-pinning-<old> nvidia-kernel-open-dkms nvidia-driver nvidia-open\napt install nvidia-driver-pinning-<new>\napt -V install nvidia-open\n```\n\nThis is what happened during initial setup: `590` failed this way against\n\nthe `7.0.x-pve` kernel branch; `615` built clean. If a future PVE kernel\n\nupgrade breaks the build again, check for a newer `nvidia-driver-pinning-*`\n\nbranch before assuming it’s a headers problem — this class of failure\n\ndoesn’t affect the A2’s hardware support (NVIDIA’s data-center driver\n\nbranches carry long backward GPU support), it’s purely about the driver\n\nsource’s kernel-API coverage.\n\n## \n\n```\nnvidia-smi\nls /dev/nvidia*\n```\n\nExpect `nvidia0`, `nvidiactl`, `nvidia-uvm`, and `nvidia-uvm-tools`. If\n\n`nvidia-uvm*` is missing (needed for CUDA unified memory), load it manually\n\nonce and add a systemd unit or udev rule to load it at boot:\n\n```\nnvidia-modprobe -c0 -u\n```\n\n## \n\nWhen creating the container (or editing an existing one):\n\n- **Options → Privileged** — uncheck “Unprivileged container” at creation\n time, since this isn’t changeable after the fact without recreating the\ncontainer.\n- Privileged avoids the UID/GID mapping headaches that unprivileged\n containers hit with GPU device ownership, and lines up with wanting NFS\n access from inside the container for the Hermes/Judge agent’s home\n directory — Kerberized or ID-mapped NFS mounts are simpler to get right\n from a privileged container than punching them through an unprivileged\ncontainer’s user namespace.\n\n## \n\n1. \nSelect the container → **Resources → Add → Device Passthrough** .\n2. \nAdd one entry per device node, typing the path directly (same four paths\non every node, since each node’s A2 enumerates the same way):\n \n\n```\n/dev/nvidia0\n/dev/nvidiactl\n/dev/nvidia-uvm\n/dev/nvidia-uvm-tools\n```\n\n3. \nStart (or restart) the container.\n\n## \n\nSame pinned driver version as the host, but install only the userspace\n\nmetapackage — **not** `nvidia-open` — since the container shares the host’s\n\nalready-loaded kernel module through the passed-through device nodes and\n\nhas no kernel headers to build a DKMS module against anyway.\n\nThe right package is `nvidia-driver-cuda` (confirmed via `dpkg -S $(which nvidia-smi)` on the host — that’s what actually owns the `nvidia-smi`\n\nbinary). Its dependency list is pure userspace (`libnvidia-ml1`,\n\n`libcuda1`-adjacent libraries, `nvidia-opencl-icd`, `nvidia-persistenced`,\n\netc.) with **no** `nvidia-driver`/DKMS dependency — unlike `nvidia-open`,\n\nwhich pulls in `nvidia-kernel-open-dkms` and will try (uselessly) to build\n\na kernel module inside the container.\n\nUsing the same apt-based repo setup as step 1 (headers/contrib/non-free\n\nsteps can be skipped inside the container):\n\n```\nwget https://developer.download.nvidia.com/compute/cuda/repos/debian13/x86_64/cuda-keyring_1.1-1_all.deb\ndpkg -i cuda-keyring_1.1-1_all.deb\napt update\napt install nvidia-driver-pinning-615\napt install nvidia-driver-cuda cuda-toolkit-13-4\n```\n\nYou may see a message during `cuda-toolkit` setup about installing\n\n`linux-headers-<kernel>` — that’s the optional GPUDirect Storage\n\n(`nvidia-fs`) component offering an unrelated DKMS module. Harmless to\n\nignore unless you specifically need GPUDirect Storage; it doesn’t affect\n\n`nvidia-smi` or CUDA inference workloads.\n\n`nvidia-driver-cuda` pulls in `nvidia-smi` and the userspace driver\n\nlibraries without touching the kernel module, since `/dev/nvidia*` is already\n\npassed through from the host.\n\n## \n\n```\nnvidia-smi\n```\n\ninside the container should report the A2. If you get a driver\n\nversion-mismatch error, the container’s userspace driver doesn’t match\n\nwhichever host it’s currently scheduled on — check that first before\n\ndigging into device permissions.\n\n# \n\nTwo things drift out from under this setup over time: the **PVE kernel**\n\n(new headers needed, DKMS has to rebuild against it) and the **driver\nversion** (host and every container’s userspace must stay in lock-step).\n\nHandle them in that order — kernel first, then confirm the module actually\n\nrebuilt — every time you run `apt upgrade` on a node.\n\n## `apt upgrade` on a host\n\n1. \n**Before upgrading** , note the currently working kernel and driver:\n \n\n```\nuname -r\nnvidia-smi   # top-right shows driver version, e.g. 615.xx\n```\n\n2. \nRun the upgrade as normal: \n\n```\napt update && apt upgrade\n```\n\n If this pulls in a new `proxmox-kernel-*` package, apt will**not**\n automatically install matching headers for it — that’s a separate\npackage you still have to install yourself (step 4).\n3. \n**Don’t reboot yet.** Check whether a new kernel was actually installed:\n \n\n```\nproxmox-boot-tool kernel list\n```\n\n Compare against `uname -r` . If a newer kernel now appears in the list,\nthat’s what you’ll boot into next.\n4. \nInstall headers for the **new** kernel*before* rebooting into it, using\nthe same pattern as initial setup:\n \n\n```\napt install proxmox-headers-$(ls /boot/vmlinuz-* | sort -V | tail -1 | sed -E 's#.*vmlinuz-([0-9]+\\.[0-9]+).*#\\1#')\n```\n\n This derives the branch from the newest kernel file in `/boot` rather\n than the currently*running* one, since after an upgrade the two\n differ. If that’s ever unclear, just read the version off\n `proxmox-boot-tool kernel list` and install`proxmox-headers-<X.Y>`\ndirectly — more explicit, harder to get wrong.\n5. \nReboot: \n\n```\nreboot\n```\n\n6. \n**After reboot, verify DKMS actually rebuilt against the new kernel**\nbefore assuming the GPU is fine:\n \n\n```\nuname -r\ndkms status\nnvidia-smi\n```\n\n `dkms status` should show the`nvidia` module built against the kernel\n `uname -r` just reported. If`nvidia-smi` fails with “NVIDIA driver\n not loaded” or similar, the headers either weren’t present at boot time\n or didn’t match — install the correct`proxmox-headers-<X.Y>` package\nand force a rebuild:\n \n\n```\ndkms autoinstall\n```\n\n If `dkms autoinstall` (or the original DKMS build during`apt install` )\n fails with compiler errors instead of a missing-headers error, that’s a\n different problem — see “Troubleshooting: DKMS build fails with\n `page_free` /`zone_device_page_init` errors” under step 1, which covers\n the kernel-API-incompatibility case actually hit during initial setup\non this cluster.\n\n## \n\nDo this deliberately, not as a side effect of `apt upgrade` — the pinning\n\npackage is what prevents an ordinary upgrade from silently jumping driver\n\nversions on you.\n\n1. \nOn **one** host, check what’s available and pick the new pinned driver\nand toolkit versions:\n \n\n```\napt list -a nvidia-driver-pinning-*\napt-cache search cuda-toolkit\n```\n\n2. \nOn **every node** , in the same session/day so nothing runs mismatched\nfor long:\n \n\n```\napt install nvidia-driver-pinning-<new-version>\napt -V install nvidia-open cuda-toolkit-<new-version>\nreboot\n```\n\n3. \nOn **every privileged LXC** using the GPU, matching the same versions:\n \n\n```\napt install nvidia-driver-pinning-<new-version>\napt install nvidia-driver-cuda cuda-toolkit-<new-version>\n```\n\n Container restart is usually enough (no reboot needed — no kernel\nmodule inside the container).\n4. \nVerify everywhere: \n\n```\nnvidia-smi   # on each host, and inside each container\n```\n\n All of them should report the same driver version. A mismatch here is\n the single most common failure mode with this setup — if any\n `nvidia-smi` errors out with a version-mismatch message, that node or\ncontainer is the one still holding the old pin.\n\n## \n\n-  Snapshot/note current `uname -r` and driver version on each node\n-  After upgrade, install headers for the new kernel *before* reboot\n-  Reboot, then confirm `dkms status` shows a build against the new\n kernel and`nvidia-smi` succeeds\n-  Only touch the driver version (`nvidia-driver-pinning-*` ) as its own\n deliberate step across all three nodes + containers together — never\nlet it drift as a side effect of a routine kernel/security upgrade", "url": "https://wpnews.pro/news/nvidia-drivers-on-debian-13-trixie", "canonical_source": "https://forum.level1techs.com/t/nvidia-drivers-on-debian-13-trixie/252574#post_13", "published_at": "2026-10-02 14:33:12+00:00", "updated_at": "2026-10-02 14:37:19.491351+00:00", "lang": "en", "topics": ["ai-infrastructure", "mlops"], "entities": ["NVIDIA", "NVIDIA Tesla A2", "Proxmox VE", "Debian 13 Trixie", "CUDA", "Claude", "cuda-keyring", "nvidia-open"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/nvidia-drivers-on-debian-13-trixie", "markdown": "https://wpnews.pro/news/nvidia-drivers-on-debian-13-trixie.md", "text": "https://wpnews.pro/news/nvidia-drivers-on-debian-13-trixie.txt", "jsonld": "https://wpnews.pro/news/nvidia-drivers-on-debian-13-trixie.jsonld"}}