RE: [PATCH 00/10] QEMU: CXL Type-2 device passthrough via vfio-pci
Manish Honap <[email protected]>
| Newsgroups | org.nongnu.qemu-devel,org.nongnu.qemu-arm |
|---|---|
| Message-ID | <IA1PR12MB90307AE3C151A0BFA5ABF998BDDA2@IA1PR12MB9030.namprd12.prod.outlook.com> |
> -----Original Message----- > From: Cédric Le Goater <[email protected]> > Sent: 13 August 2026 23:13 > To: Manish Honap <[email protected]>; [email protected]; Ankit Agrawal > <[email protected]>; [email protected]; [email protected]; > [email protected]; Srirangan Madhavan <[email protected]>; > [email protected]; [email protected]; [email protected]; > [email protected]; [email protected]; [email protected]; > [email protected]; [email protected]; [email protected] > Cc: Krishnakant Jaju <[email protected]>; Vikram Sethi <[email protected]>; > Zhi Wang <[email protected]>; [email protected]; [email protected] > Subject: Re: [PATCH 00/10] QEMU: CXL Type-2 device passthrough via vfio- > pci > > External email: Use caution opening links or attachments > > > Hello Manish, > > On 8/13/26 15:06, [email protected] wrote: > > From: Manish Honap <[email protected]> > > > > This series adds QEMU support for passing a CXL Type-2 device (an > > accelerator with host-managed device memory, e.g. a GPU) to a guest > > via vfio-pci. The guest drives its own virtual endpoint HDM decoder > > and QEMU maps the device memory at the guest physical address the > > guest commits, while the host owns the host physical placement. > > > > It is a full rewrite of the RFC v1 series [1] on clean upstream > > master, reworked to address Jonathan's review. > > > > Base: master (post pull-9p-20260725), commit 6333226c2a. > > > > > > Kernel dependency > > ----------------- > > > > Pairs with the kernel vfio-cxl series "vfio/cxl: CXL Type-2 device > > passthrough" [2]. > > - The kernel exposes the device memory as an HPA-backed VFIO region, > traps > > the HDM decoder block and runs its lock-on-commit FSM, handles the > CXL > > DVSEC (including guest-triggered reset), and reports two things > through > > VFIO: > > - A device flag (VFIO_DEVICE_FLAGS_CXL) and > > - The component-register geometry > (VFIO_REGION_INFO_CAP_CXL_COMP_REGS). > > > > > > Sample supported topology > > ------------------------- > > > > Guest disk, network, and system-RAM lines are omitted: > > > > -machine virt,accel=kvm,gic-version=3,hmat=on,cxl=on,ras=on, \ > > highmem-mmio-size=4T > > -object iommufd,id=iommufd0 > > -device pxb-cxl,bus_nr=12,bus=pcie.0,id=cxl.1 > > -device cxl-rp,port=1,bus=cxl.1,id=rport0.1,chassis=4, \ > > pref64-reserve=2G,mem-reserve=1G > > -M cxl-fmw.0.targets.0=cxl.1,cxl-fmw.0.size=256G > > -device arm-smmuv3,primary-bus=cxl.1,id=smmuv3.0,accel=on,ats=on, \ > > ril=on,ssidsize=8,oas=48 > > -device vfio-pci-nohotplug,host=<BDF>,bus=rport0.1,id=dev0, \ > > iommufd=iommufd0 > > -object acpi-generic-initiator,id=gi0,pci-dev=dev0,node=2 > > ... (one acpi-generic-initiator per guest NUMA node the HDM memory > > backs) > > > > > > Address model > > ------------- > > > > The kernel fixes device memory to a host physical range before the > > guest sees the device, and hardware presents a firmware-committed, > > locked endpoint HDM decoder whose registers hold that host physical > base. > > > > QEMU never exposes Host physical base. It virtualizes the decoder base > > registers in the trapped component-register read and returns the base > > of the device's CFMWS window, a guest physical address, so the guest > only ever sees a GPA. > > > > QEMU maps the RAM-device region (backed by the fixed HPA) at that CFMWS > base. > > > > The kernel never learns the GPA and the guest never learns the HPA. > > > > Because the decoder is already committed at boot, no guest commit > > write triggers the mapping. QEMU maps once the guest enables memory > > decoding (the Command register Memory-Space bit) and re-checks on any > > decoder control write, so the region enters the guest address space > > and the IOAS while the device is live. > > > > When the guest clears Memory-Space, QEMU withdraws the mapping, > > matching the kernel's revoke of the backing PTEs on the same write, so > > a guest access during the disabled interval cannot fault a zapped > > mapping and stop the VM; the enable path re-installs it. > > > > This also helps to keep the guest-visible base as the CFMWS base by > > construction, and the host physical placement stays with the kernel. > > > > The RFC added the pxb-cxl _DSM because OS may treat PCI configuration > > as reassignable and a BAR move would break the CXL.mem mapping. This > > change keeps it narrower than the RFC v1 _DSM: it applies only to the > > host bridge that carries the passed-through CXL device. > > > > > > Reset > > ----- > > > > There is no QEMU reset patch. A guest CXL reset is a DVSEC write that > > lands in vfio config space and is handled entirely by the host kernel, > > which stamps the outcome into DVSEC STATUS2. The kernel re-commits > > this firmware-fixed decoder across the reset, and the guest reaches > > its memory through the mapping QEMU already installed. > > > > > > Reviewer feedback addressed > > --------------------------- > > > > The RFC [1] thread has the full discussion. > > > > Jonathan Cameron > > - The high-MMIO window and the cxl-fmws-base property are removed. > The > > guest-visible base is the CFMWS base by construction, so it equals > the > > decoder base and stays stable without a hack. > > - The guest programs its virtual (GPA) decoder and QEMU maps the > memory at > > commit time. > > - The host owns the HPA, resolved before the guest sees the device. > > - The committed decoder is the fast path this series ships. The > > guest-programmed uncommitted case is delivered by cxl-core > resolving the > > range at enumeration, not by a QEMU or vfio dynamic branch, so it > is a > > later cxl-core item this series does not depend on. > > - FIRMWARE_COMMITTED is dropped; the cap is no longer exposed. > > - The one-endpoint, non-interleaved, no-switch topology is enforced > at > > realize. > > - PCI/BAR configuration and CXL.mem stay independent. > > > > > > Validation > > ---------- > > > > - Every patch passes scripts/checkpatch.pl --codespell --strict > > (patch 1 carries the expected imported-from-Linux warning) > > - The series applies cleanly on the stated base. > > - Ran couple of rounds of masoncl/review-prompts on this series before > posting. > > > > > > Pending items > > ------------- > > > > Future enhancements for the multi-decoder, interleaved devices and > > switched topologies. > > > > Trapped CXL RAS registers are planned as a new VFIO region subtype > > that the same region-by-subtype detection already handles, not as a > > change to the component-register cap. > > > > The bios-tables test refresh for the new _DSM is still to be added. > > Before diving into CXL and doing a review of the QEMU patches, could you > provide a status update on the kernel vfio-cxl series [2] ? > > The QEMU side is a consumer of the VFIO regions and capabilities the > kernel exposes, and patch 1 notes the capability ID is still provisional. > Knowing where the kernel series stands would help. That said, the UAPI > looks simple enough. > > What has become of the CXL Type-2 device emulation effort from Zhi Wang? > > https://lore.kernel.org/qemu-devel/20241212130422.69380-1- > [email protected]/ > > That series introduced a bare-minimum emulated CXL Type-2 device in QEMU, > so kernel and QEMU developers could test the CXL Type-2 stack without real > hardware. I liked the idea quite a lot. > > CXL is still new and complex, and hardware is scarce. Having an emulated > device and passthrough support gives us a full environment to flush out > design issues across the stack (Linux kernel, QEMU, libvirt) and in > associated subsystems like IOMMU and ACPI, without needing physical > devices. It is also useful for education and for running CI tests. Do you > see both efforts as complementary, and is the emulation work still active? > > I will take a closer look at the series after some time off. > Hello Cédric, Thank you for your suggestion. Kernel status: Current posting of the vfio-cxl series takes the maintainer comments and refactors the code to what CXL core requires. The series adds a load-on-demand vfio-cxl provider to vfio-pci-core for a bound CXL Type-2 accelerator. The module exclusively claims the firmware-committed HDM host-physical range and exposes that coherent memory as an mmap-only VFIO region. DVSEC and HDM decoder registers are copied into a per-open shadow so the guest can program them without touching the host register space. The two UAPI values (FLAGS_CXL bit, COMP_REGS cap id) are settled for this posting. I will make sure to include the kernel status in the QEMU series cover letter from next time. I agree on emulated device support. Zhi's series had two parts: a bare-minimum emulated cxl-accel device, and an early passthrough stub. This series rewrites the passthrough side after Jonathan's review. I want a bit more time on your suggestion and will get back with how we can fold the emulated-device work in. > Thanks, > > > C. > > > > > >