RE: [PATCH 00/10] QEMU: CXL Type-2 device passthrough via vfio-pci

Manish Honap <[email protected]>
Newsgroups org.nongnu.qemu-devel,org.nongnu.qemu-arm
Message-ID <IA1PR12MB90307AE3C151A0BFA5ABF998BDDA2@IA1PR12MB9030.namprd12.prod.outlook.com>

> -----Original Message-----
> From: Cédric Le Goater <[email protected]>
> Sent: 13 August 2026 23:13
> To: Manish Honap <[email protected]>; [email protected]; Ankit Agrawal
> <[email protected]>; [email protected]; [email protected];
> [email protected]; Srirangan Madhavan <[email protected]>;
> [email protected]; [email protected]; [email protected];
> [email protected]; [email protected]; [email protected];
> [email protected]; [email protected]; [email protected]
> Cc: Krishnakant Jaju <[email protected]>; Vikram Sethi <[email protected]>;
> Zhi Wang <[email protected]>; [email protected]; [email protected]
> Subject: Re: [PATCH 00/10] QEMU: CXL Type-2 device passthrough via vfio-
> pci
> 
> External email: Use caution opening links or attachments
> 
> 
> Hello Manish,
> 
> On 8/13/26 15:06, [email protected] wrote:
> > From: Manish Honap <[email protected]>
> >
> > This series adds QEMU support for passing a CXL Type-2 device (an
> > accelerator with host-managed device memory, e.g. a GPU) to a guest
> > via vfio-pci. The guest drives its own virtual endpoint HDM decoder
> > and QEMU maps the device memory at the guest physical address the
> > guest commits, while the host owns the host physical placement.
> >
> > It is a full rewrite of the RFC v1 series [1] on clean upstream
> > master, reworked to address Jonathan's review.
> >
> > Base: master (post pull-9p-20260725), commit 6333226c2a.
> >
> >
> > Kernel dependency
> > -----------------
> >
> > Pairs with the kernel vfio-cxl series "vfio/cxl: CXL Type-2 device
> > passthrough" [2].
> > - The kernel exposes the device memory as an HPA-backed VFIO region,
> traps
> >    the HDM decoder block and runs its lock-on-commit FSM, handles the
> CXL
> >    DVSEC (including guest-triggered reset), and reports two things
> through
> >    VFIO:
> >    - A device flag (VFIO_DEVICE_FLAGS_CXL) and
> >    - The component-register geometry
> (VFIO_REGION_INFO_CAP_CXL_COMP_REGS).
> >
> >
> > Sample supported topology
> > -------------------------
> >
> > Guest disk, network, and system-RAM lines are omitted:
> >
> >    -machine virt,accel=kvm,gic-version=3,hmat=on,cxl=on,ras=on, \
> >             highmem-mmio-size=4T
> >    -object iommufd,id=iommufd0
> >    -device pxb-cxl,bus_nr=12,bus=pcie.0,id=cxl.1
> >    -device cxl-rp,port=1,bus=cxl.1,id=rport0.1,chassis=4, \
> >            pref64-reserve=2G,mem-reserve=1G
> >    -M cxl-fmw.0.targets.0=cxl.1,cxl-fmw.0.size=256G
> >    -device arm-smmuv3,primary-bus=cxl.1,id=smmuv3.0,accel=on,ats=on, \
> >            ril=on,ssidsize=8,oas=48
> >    -device vfio-pci-nohotplug,host=<BDF>,bus=rport0.1,id=dev0, \
> >            iommufd=iommufd0
> >    -object acpi-generic-initiator,id=gi0,pci-dev=dev0,node=2
> >    ... (one acpi-generic-initiator per guest NUMA node the HDM memory
> > backs)
> >
> >
> > Address model
> > -------------
> >
> > The kernel fixes device memory to a host physical range before the
> > guest sees the device, and hardware presents a firmware-committed,
> > locked endpoint HDM decoder whose registers hold that host physical
> base.
> >
> > QEMU never exposes Host physical base. It virtualizes the decoder base
> > registers in the trapped component-register read and returns the base
> > of the device's CFMWS window, a guest physical address, so the guest
> only ever sees a GPA.
> >
> > QEMU maps the RAM-device region (backed by the fixed HPA) at that CFMWS
> base.
> >
> > The kernel never learns the GPA and the guest never learns the HPA.
> >
> > Because the decoder is already committed at boot, no guest commit
> > write triggers the mapping. QEMU maps once the guest enables memory
> > decoding (the Command register Memory-Space bit) and re-checks on any
> > decoder control write, so the region enters the guest address space
> > and the IOAS while the device is live.
> >
> > When the guest clears Memory-Space, QEMU withdraws the mapping,
> > matching the kernel's revoke of the backing PTEs on the same write, so
> > a guest access during the disabled interval cannot fault a zapped
> > mapping and stop the VM; the enable path re-installs it.
> >
> > This also helps to keep the guest-visible base as the CFMWS base by
> > construction, and the host physical placement stays with the kernel.
> >
> > The RFC added the pxb-cxl _DSM because OS may treat PCI configuration
> > as reassignable and a BAR move would break the CXL.mem mapping. This
> > change keeps it narrower than the RFC v1 _DSM: it applies only to the
> > host bridge that carries the passed-through CXL device.
> >
> >
> > Reset
> > -----
> >
> > There is no QEMU reset patch. A guest CXL reset is a DVSEC write that
> > lands in vfio config space and is handled entirely by the host kernel,
> > which stamps the outcome into DVSEC STATUS2. The kernel re-commits
> > this firmware-fixed decoder across the reset, and the guest reaches
> > its memory through the mapping QEMU already installed.
> >
> >
> > Reviewer feedback addressed
> > ---------------------------
> >
> > The RFC [1] thread has the full discussion.
> >
> > Jonathan Cameron
> >    - The high-MMIO window and the cxl-fmws-base property are removed.
> The
> >      guest-visible base is the CFMWS base by construction, so it equals
> the
> >      decoder base and stays stable without a hack.
> >    - The guest programs its virtual (GPA) decoder and QEMU maps the
> memory at
> >      commit time.
> >    - The host owns the HPA, resolved before the guest sees the device.
> >    - The committed decoder is the fast path this series ships. The
> >      guest-programmed uncommitted case is delivered by cxl-core
> resolving the
> >      range at enumeration, not by a QEMU or vfio dynamic branch, so it
> is a
> >      later cxl-core item this series does not depend on.
> >    - FIRMWARE_COMMITTED is dropped; the cap is no longer exposed.
> >    - The one-endpoint, non-interleaved, no-switch topology is enforced
> at
> >      realize.
> >    - PCI/BAR configuration and CXL.mem stay independent.
> >
> >
> > Validation
> > ----------
> >
> > - Every patch passes scripts/checkpatch.pl --codespell --strict
> >    (patch 1 carries the expected imported-from-Linux warning)
> > - The series applies cleanly on the stated base.
> > - Ran couple of rounds of masoncl/review-prompts on this series before
> posting.
> >
> >
> > Pending items
> > -------------
> >
> > Future enhancements for the multi-decoder, interleaved devices and
> > switched topologies.
> >
> > Trapped CXL RAS registers are planned as a new VFIO region subtype
> > that the same region-by-subtype detection already handles, not as a
> > change to the component-register cap.
> >
> > The bios-tables test refresh for the new _DSM is still to be added.
> 
> Before diving into CXL and doing a review of the QEMU patches, could you
> provide a status update on the kernel vfio-cxl series [2] ?
> 
> The QEMU side is a consumer of the VFIO regions and capabilities the
> kernel exposes, and patch 1 notes the capability ID is still provisional.
> Knowing where the kernel series stands would help. That said, the UAPI
> looks simple enough.
> 
> What has become of the CXL Type-2 device emulation effort from Zhi Wang?
> 
>    https://lore.kernel.org/qemu-devel/20241212130422.69380-1-
> [email protected]/
> 
> That series introduced a bare-minimum emulated CXL Type-2 device in QEMU,
> so kernel and QEMU developers could test the CXL Type-2 stack without real
> hardware. I liked the idea quite a lot.
> 
> CXL is still new and complex, and hardware is scarce. Having an emulated
> device and passthrough support gives us a full environment to flush out
> design issues across the stack (Linux kernel, QEMU, libvirt) and in
> associated subsystems like IOMMU and ACPI, without needing physical
> devices. It is also useful for education and for running CI tests. Do you
> see both efforts as complementary, and is the emulation work still active?
> 
> I will take a closer look at the series after some time off.
> 

Hello Cédric,

Thank you for your suggestion.

Kernel status: Current posting of the vfio-cxl series takes the maintainer
comments and refactors the code to what CXL core requires. The series adds
a load-on-demand vfio-cxl provider to vfio-pci-core for a bound CXL Type-2
accelerator. The module exclusively claims the firmware-committed HDM
host-physical range and exposes that coherent memory as an mmap-only VFIO
region. DVSEC and HDM decoder registers are copied into a per-open shadow
so the guest can program them without touching the host register space.

The two UAPI values (FLAGS_CXL bit, COMP_REGS cap id) are settled for
this posting.

I will make sure to include the kernel status in the QEMU series cover
letter from next time.

I agree on emulated device support. Zhi's series had two parts: a
bare-minimum emulated cxl-accel device, and an early passthrough stub.
This series rewrites the passthrough side after Jonathan's review. I
want a bit more time on your suggestion and will get back with how we
can fold the emulated-device work in.

> Thanks,
> 
> 
> C.
> 
> 
> 
> 
> 
>
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.