RE: [RFCv2 PATCH 0/6] Support memory hotplug/unplug for TDX CoCo guests

"Duan, Zhenzhong" <[email protected]>
Newsgroups dev.linux.lists.virtualization,dev.linux.lists.linux-coco,org.kernel.vger.linux-kernel
Message-ID <IA3PR11MB91365F3F771C43C4BA7E61B392AD2@IA3PR11MB9136.namprd11.prod.outlook.com>

>-----Original Message-----
>From: David Hildenbrand (Arm) <[email protected]>
>Subject: Re: [RFCv2 PATCH 0/6] Support memory hotplug/unplug for TDX CoCo
>guests
>
>On 6/23/26 12:17, Zhenzhong Duan wrote:
>> This RFCv2 series implements comprehensive support for virtio-mem and ACPI
>> DIMM memory hotplug/unplug in Intel TDX confidential computing guests.
>> It explores the start-private memory approach utilizing the native
>> TDG.MEM.PAGE.RELEASE API.
>>
>> We are seeking feedback from Kiryl on the CoCo guest implementation, MM
>> experts on DIMM & virio-mem memory hotplug integration and broader
>> virtio/CoCo community input on the overall approach. We are not seeking
>> x86 maintainer review at this stage.
>>
>> == Changes from RFC v1 ==
>>
>> - Eliminated callback infrastructure: Dropped plug callback and replaced
>>   unplug callback with platform-level unaccept function into core MM
>>   hotplug and virtio-mem subsystems.
>> - Added comprehensive bitmap tracking: Introduced a "plugged" bitmap
>>   alongside the unaccepted bitmap to track populated hotplug memory
>>   states to support load_unaligned_zeropad().
>> - Enhanced SRAT parsing: Extended the EFI stub to parse ACPI SRAT tables
>>   early, ensuring hotpluggable ranges are tracked from initial boot.
>>
>> For more introduction about the background or other efforts in community,
>> please check the RFCv1 cover letter [1].
>>
>> == Technical Approach ==
>>
>> - Early SRAT Integration: A lightweight EFI stub parser scans ACPI SRAT
>>   tables to identify hotpluggable ranges and adjust bitmap boundaries
>>   early, avoiding the overhead of the full ACPI subsystem.
>> - Comprehensive Bitmap Tracking: Introduces a "plugged" bitmap right
>>   after the unaccepted bitmap. Both static and hotplugged memory are
>>   tracked, allowing the guest to map which ranges are populated by the
>>   VMM. This prevents acceptance beyond plugged memory boundaries due to
>>   load_unaligned_zeropad() operations.
>> - Platform Extensibility: Exposes generic CoCo memory interfaces. Other
>>   confidential platforms (like AMD SEV-SNP) can easily adopt this by
>>   hooking their specific mechanisms into arch_unaccept_memory().
>> - Hotplug & Guest Control: Integrates platform-level unaccept logic
>>   into ACPI hotplug and virtio-mem handlers. Uses TDG.MEM.PAGE.RELEASE
>>   for TDX to explicitly set memory to the "unaccepted" state during
>>   unplug, removing host hole-punching dependencies.
>> - Kexec Handover: Leverages existing EFI mechanisms to seamlessly hand
>>   over both the extended unaccepted bitmap and the new plugged bitmap
>>   across kexec boundaries.
>>
>> == Testing ==
>>
>> - dimm and virtio-mem memory hotplug/unplug
>> - lazy and eager accept
>> - kexec/kdump with hotplugged memory
>>
>> This is tested with Marc-André Lureau's newest qemu series [2]
>
>What's the status of this?

Marc's QEMU series is merged.
For this series, following feedback from Kirill and Pratik, the preferred approach
is updating the UEFI spec for hotplug memory ranges rather than parsing SRAT
at the EFI stage. Pratik is already pushing this forward, I am currently waiting on
his RFCs. If he hasn't taken over the entire implementation, I can rebase my
remaining patches on top of his work.

Hi Pratik, have you sent your UEFI RFC out yet? Just wanted to make sure
I didn't miss your thread.

>
>I am still not sure whether we shouldn't perform acceptance from
>move_pfn_range_to_zone() and from memory notifiers / generic_online_page.

My understanding is that we already have full support for lazy and eager acceptance
in generic_online_page() for static memory. We should be able to reuse that for
hotplug memory and avoid adding acceptance logic in other places.

All we need is extending unaccept_bitmap and adding new plugged_bitmap to support
hotplug memory. I updated accept_memory() to check both bitmaps to determine
which memory should be accepted.

>
>In particular, it's unclear to me how virtio-mem (which uses interfaces to
>add/remove memory) interacts with unaccept_memory / coco bitmap.

It works the same way as a physical DIMM: when memory is plugged, the
corresponding bits in plugged_bitmap are set, and vice versa.

Memory acceptance is already handled in generic_online_page(), so we 
do not need to do it inside virtio-mem. However, we do need to call
unaccept_memory() during a memory unplug event in virtio-mem.

Currently, tdx_unaccept_memory() can act as a no-op since QEMU handles 
hole-punching the private memory. That said, we still need to invoke 
unaccept_memory() to properly update the unaccept_bitmap bits.

>
>Can we have an overall design view on what happens at which stage when adding
>/
>removing memory through virtio-mem?

I have put together a design view summary for virtio-mem below.
Please let me know if this looks correct or if we should adjust the framing.

Design Overview
---------------
We maintain system stability and state safety using two metadata tracking
layers during dynamic memory resizing operations:
1. plugged_bitmap: Explicitly tracks blocks plugged into the guest.
   This protects load_unaligned_zeropad() from reading omitted memory
   holes, preventing catastrophic guest crashes.
2. unaccept_bitmap: Explicitly tracks the secure page initialization state.

Step-by-Step Lifecycle Stages
-----------------------------
Using sub-block hotplug of a new memory block with eager acceptance
as an example:

1. Memory Addition (Plug) Stage
   a. Host notifies guest -> virtio-mem driver handles the plug event.
   b. Driver marks the allocated memory ranges in 'plugged_bitmap' by
      calling coco_set_plugged_bitmap(addr, size, true). During this stage,
      a plug request is also sent to the VMM.
   c. Driver adds memory blocks via add_memory_resource().
   d. Subsystem triggers online -> driver callback runs generic_online_page().
   e. generic_online_page() cross-references 'plugged_bitmap' and
      'unaccept_bitmap' to execute secure page acceptance (e.g.,
      TDG.MEM.PAGE.ACCEPT).
   f. Freshly accepted pages are freed directly into the buddy allocator.

2. Memory Removal (Unplug) Stage
   a. Host requests memory reclamation -> virtio-mem targets the range.
   b. Pages are transitioned to a fake-offline state via virtio_mem_fake_offline().
   c. Driver calls unaccept_memory() to mark pages back as unaccepted in
      'unaccept_bitmap', and calls coco_set_plugged_bitmap(addr, size, false)
      to mark pages unplugged in 'plugged_bitmap'. During this stage, an unplug
      request is also sent to the VMM to hole-punch the backing private memory.
   d. Driver offlines and removes the memory blocks from the kernel via
      offline_and_remove_memory().

BRs,
Zhenzhong
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.