Re: [PATCH 00/10] QEMU: CXL Type-2 device passthrough via vfio-pci

Cédric Le Goater <[email protected]>
Newsgroups org.nongnu.qemu-arm,org.nongnu.qemu-devel
Message-ID <[email protected]>
Hello Manish,

On 8/13/26 15:06, [email protected] wrote:
> From: Manish Honap <[email protected]>
> 
> This series adds QEMU support for passing a CXL Type-2 device (an
> accelerator with host-managed device memory, e.g. a GPU) to a guest via
> vfio-pci. The guest drives its own virtual endpoint HDM decoder and QEMU
> maps the device memory at the guest physical address the guest commits,
> while the host owns the host physical placement.
> 
> It is a full rewrite of the RFC v1 series [1] on clean upstream master,
> reworked to address Jonathan's review.
> 
> Base: master (post pull-9p-20260725), commit 6333226c2a.
> 
> 
> Kernel dependency
> -----------------
> 
> Pairs with the kernel vfio-cxl series "vfio/cxl: CXL Type-2 device
> passthrough" [2].
> - The kernel exposes the device memory as an HPA-backed VFIO region, traps
>    the HDM decoder block and runs its lock-on-commit FSM, handles the CXL
>    DVSEC (including guest-triggered reset), and reports two things through
>    VFIO:
>    - A device flag (VFIO_DEVICE_FLAGS_CXL) and
>    - The component-register geometry (VFIO_REGION_INFO_CAP_CXL_COMP_REGS).
> 
> 
> Sample supported topology
> -------------------------
> 
> Guest disk, network, and system-RAM lines are omitted:
> 
>    -machine virt,accel=kvm,gic-version=3,hmat=on,cxl=on,ras=on, \
>             highmem-mmio-size=4T
>    -object iommufd,id=iommufd0
>    -device pxb-cxl,bus_nr=12,bus=pcie.0,id=cxl.1
>    -device cxl-rp,port=1,bus=cxl.1,id=rport0.1,chassis=4, \
>            pref64-reserve=2G,mem-reserve=1G
>    -M cxl-fmw.0.targets.0=cxl.1,cxl-fmw.0.size=256G
>    -device arm-smmuv3,primary-bus=cxl.1,id=smmuv3.0,accel=on,ats=on, \
>            ril=on,ssidsize=8,oas=48
>    -device vfio-pci-nohotplug,host=<BDF>,bus=rport0.1,id=dev0, \
>            iommufd=iommufd0
>    -object acpi-generic-initiator,id=gi0,pci-dev=dev0,node=2
>    ... (one acpi-generic-initiator per guest NUMA node the HDM memory backs)
> 
> 
> Address model
> -------------
> 
> The kernel fixes device memory to a host physical range before the guest
> sees the device, and hardware presents a firmware-committed, locked
> endpoint HDM decoder whose registers hold that host physical base.
> 
> QEMU never exposes Host physical base. It virtualizes the decoder base registers
> in the trapped component-register read and returns the base of the device's
> CFMWS window, a guest physical address, so the guest only ever sees a GPA.
> 
> QEMU maps the RAM-device region (backed by the fixed HPA) at that CFMWS base.
> 
> The kernel never learns the GPA and the guest never learns the HPA.
> 
> Because the decoder is already committed at boot, no guest commit write
> triggers the mapping. QEMU maps once the guest enables memory decoding
> (the Command register Memory-Space bit) and re-checks on any decoder
> control write, so the region enters the guest address space and the IOAS
> while the device is live.
> 
> When the guest clears Memory-Space, QEMU withdraws the mapping, matching the
> kernel's revoke of the backing PTEs on the same write, so a guest access during
> the disabled interval cannot fault a zapped mapping and stop the VM; the enable
> path re-installs it.
> 
> This also helps to keep the guest-visible base as the CFMWS base by
> construction, and the host physical placement stays with the kernel.
> 
> The RFC added the pxb-cxl _DSM because OS may treat PCI configuration as
> reassignable and a BAR move would break the CXL.mem mapping. This change keeps
> it narrower than the RFC v1 _DSM: it applies only to the host bridge that
> carries the passed-through CXL device.
> 
> 
> Reset
> -----
> 
> There is no QEMU reset patch. A guest CXL reset is a DVSEC write that
> lands in vfio config space and is handled entirely by the host kernel,
> which stamps the outcome into DVSEC STATUS2. The kernel re-commits this
> firmware-fixed decoder across the reset, and the guest reaches its memory
> through the mapping QEMU already installed.
> 
> 
> Reviewer feedback addressed
> ---------------------------
> 
> The RFC [1] thread has the full discussion.
> 
> Jonathan Cameron
>    - The high-MMIO window and the cxl-fmws-base property are removed. The
>      guest-visible base is the CFMWS base by construction, so it equals the
>      decoder base and stays stable without a hack.
>    - The guest programs its virtual (GPA) decoder and QEMU maps the memory at
>      commit time.
>    - The host owns the HPA, resolved before the guest sees the device.
>    - The committed decoder is the fast path this series ships. The
>      guest-programmed uncommitted case is delivered by cxl-core resolving the
>      range at enumeration, not by a QEMU or vfio dynamic branch, so it is a
>      later cxl-core item this series does not depend on.
>    - FIRMWARE_COMMITTED is dropped; the cap is no longer exposed.
>    - The one-endpoint, non-interleaved, no-switch topology is enforced at
>      realize.
>    - PCI/BAR configuration and CXL.mem stay independent.
> 
> 
> Validation
> ----------
> 
> - Every patch passes scripts/checkpatch.pl --codespell --strict
>    (patch 1 carries the expected imported-from-Linux warning)
> - The series applies cleanly on the stated base.
> - Ran couple of rounds of masoncl/review-prompts on this series before posting.
> 
> 
> Pending items
> -------------
> 
> Future enhancements for the multi-decoder, interleaved devices and switched
> topologies.
> 
> Trapped CXL RAS registers are planned as a new VFIO region subtype that the same
> region-by-subtype detection already handles, not as a change to the
> component-register cap.
> 
> The bios-tables test refresh for the new _DSM is still to be added.

Before diving into CXL and doing a review of the QEMU patches, could you
provide a status update on the kernel vfio-cxl series [2] ?

The QEMU side is a consumer of the VFIO regions and capabilities the
kernel exposes, and patch 1 notes the capability ID is still provisional.
Knowing where the kernel series stands would help. That said, the UAPI
looks simple enough.

What has become of the CXL Type-2 device emulation effort from Zhi Wang?

   https://lore.kernel.org/qemu-devel/[email protected]/

That series introduced a bare-minimum emulated CXL Type-2 device in
QEMU, so kernel and QEMU developers could test the CXL Type-2 stack
without real hardware. I liked the idea quite a lot.

CXL is still new and complex, and hardware is scarce. Having an
emulated device and passthrough support gives us a full environment to
flush out design issues across the stack (Linux kernel, QEMU, libvirt)
and in associated subsystems like IOMMU and ACPI, without needing
physical devices. It is also useful for education and for running CI
tests. Do you see both efforts as complementary, and is the emulation
work still active?

I will take a closer look at the series after some time off.

Thanks,


C.
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.