gfx1100: KFD leaks process VRAM after CP fault during AQL queue abort; device unusable until reboot

"michael wizardsound.com" <[email protected]>
Newsgroups org.freedesktop.lists.amd-gfx
Message-ID <SN7PR03MB72594F097881FCD1633872C8A2AF2@SN7PR03MB7259.namprd03.prod.outlook.com>
Hi AMD,

I submitted some findings and I wanted to send this over.  I hope this helps.

On a 7900 XTX in a Proxmox host (kernel 6.14.8-2-pve, ROCm userspaces 7.2.x/7.14 in an LXC sharing the host kernel), a userspace bug (rocm-systems#10021; root cause filed as rocm-systems#10713, runtime fix proposed in rocm-systems PR #10714) can place a malformed packet on a user AQL queue. The CP raises gfx_v11_0_bad_op_irq ("Illegal opcode in command stream") followed by an HSA_STATUS_ERROR_INVALID_PACKET_FORMAT queue abort. After the process exits, its /sys/class/kfd/kfd/proc/<pid> entry persists and ~25.7 GB VRAM stays pinned. amdgpu mode1 reset succeeds but does not release it; unbind succeeds but re-probe fails with EEXIST on mem_info_preempt_used (stale sysfs child from the wedged teardown), and subsequent per-device remove/bind attempts leave D-state tasks in device_release_driver_internal / device_driver_attach. Only a reboot recovers the device. Clean SIGSEGV deaths of the same workload never leak; it is specific to the CP-fault teardown. Full logs, timelines, and a userspace reproducer chain available; patches and standalone reproducers for the userspace side are at https://github.com/OhgurLabs/rocm-hipgraph-fixes.

Michael Levesque, MBA
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.