Re: [PATCH v4 3/6] KVM: arm64: Add auto DBM support for hardware dirty tracking

Tian Zheng <[email protected]> Mon, 3 Aug 2026 21:57:46 +0800
Newsgroups dev.linux.lists.kvmarm,org.infradead.lists.linux-arm-kernel,org.kernel.vger.kvm,org.kernel.vger.linux-kernel
Message-ID <[email protected]>

On 8/3/2026 6:21 PM, Leonardo Bras wrote:
> On Mon, Aug 03, 2026 at 12:04:24PM +0800, Tian Zheng wrote:
>>
>>
>> On 8/3/2026 9:33 AM, Tian Zheng wrote:
>>>>>>>>> 09, 2026 at 06:40:23PM +0800, Tian Zheng wrote:
>>>>>>>>>> -    if (prot & KVM_PGTABLE_PROT_W)
>>>>>>>>>> +    if (prot & KVM_PGTABLE_PROT_W) {
>>>>>>>>>>              set |= KVM_PTE_LEAF_ATTR_LO_S2_S2AP_W;
>>>>>>>>>>
>>>>>>>>>> +        /*
>>>>>>>>>> +         * No DEVICE filter needed here:
>>>>>>>>>> relax_perms is only called
>>>>>>>>>> +         * on FSC_PERM faults. Device pages
>>>>>>>>>> always get full RW from
>>>>>>>>>> +         * initial mapping and are never write-protected during
>>>>>>>>>> +         * migration, so they never trigger a permission fault.
>>>>>>>>>> +         */
>>>>>>>>>> +        if (pgt->flags & KVM_PGTABLE_S2_DBM)
>>>>>>>>>> +            set |= KVM_PTE_LEAF_ATTR_HI_S2_DBM;
>>>>>>>>>> +    } else {
>>>>>>>>>> +        /*
>>>>>>>>>> +         * Clear DBM on W→RO downgrade to prevent hardware from
>>>>>>>>>> +         * silently upgrading RO+DBM back to W+dirty, which would
>>>>>>>>>> +         * bypass KVM's write tracking and cause data corruption.
>>>>>>>>>> +         */
>>>>>>>>>> +        clr |= KVM_PTE_LEAF_ATTR_HI_S2_DBM;
>>>>>>>>>> +    }
>>>>>>>>>> +
>>>>>>>>> This block makes it pretty evident that the DBM bit really *is* the
>>>>>>>>> write permission bit. I'd much rather we
>>>>>>>>> introduce the concept of dirty
>>>>>>>>> state to the page table library and migrate the abstract write
>>>>>>>>> permission to the DBM field, even if we don't have FEAT_HAFDBS.
>>>>>>>>>
>>>>>>>
>>>>>>> Ohh, that's an amazing idea!
>>>>>>
>>>>>> Thinking about that again...
>>>>>> If we adopt the encoding with DBM being the write-permission
>>>>>> bit, and all
>>>>>> PTEs have it since the start, how can we have lazy-splitting happening?
>>>>>>
>>>>>> Only way I think of is removing both DBM and S2_S2AP_W bit
>>>>>> from writable
>>>>>> PTEs during dirty-track enable, and re-adding them during
>>>>>> the first write
>>>>>> fault. If we don't remove the DBM bit, systems with HDBSS
>>>>>> would just dirty
>>>>>> it by hardware, without causing a fault.
>>>>>>
>>>>>> DBM=0 would need to happen only in the first write-protect (only on
>>>>>> lazy-splitting). All other write-protecting would just clean
>>>>>> the S2_S2AP_W
>>>>>> bit, as everything is already split.
>>>>>>
>>>>>> Is that what was intended?
>>>>>>
>>>>>> Thanks!
>>>>>> Leo
>>>>>>
>>>>> Hi Leo,
>>>>>
>>>>> I think the cleanest way to handle this is to simply avoid setting DBM
>>>>> on block mappings. If we only set DBM on page-level PTEs, then block
>>>>> mappings will naturally stay DBM=0 and trigger a write fault on first
>>>>> access — exactly what we need for lazy splitting.
>>>>>
>>>>> When the fault occurs, the block gets split into page-level PTEs, and at
>>>>> that point we can set DBM=1 on the resulting leaf entries. This way:
>>>>>
>>>>> 1. Lazy split works naturally (fault -> split -> set DBM=1)
>>>>>
>>>>> 2. No need to clear DBM globally at dirty-track enable
>>>>>
>>>>> 3. No special handling for block mappings
>>>>>
>>>>> So I think global DBM is still viable — we just need to filter out block
>>>>> mappings when setting the DBM bit. That way the lazy split path
>>>>> is preserved
>>>>> without extra complexity.
>>>>
>>>> Hi Tian,
>>>>
>>>> Humm, but would not that be contrary to what Oliver suggested:
>>>> changing the
>>>> encoding from the PTE for all entries?
>>>>
>>>> (Like, if the PTE is writable, it has to have DBM set)
>>>>
>>>> IIUC what you said, on first faulting of the page in the VM:
>>>> - If the entry is a page (level-3 leaf) and writable, add DBM
>>>> - If it's a block entry (leaf but not a level-3), don't add DBM
>>>>
>>>> So after we enable dirty-logging:
>>>> - a level-3 entry would not fault, using HDBSS, and
>>>> - a block entry would fault, do the splitting, and add DBM to level-3
>>>>     entries during the split.
>>>>
>>>> If I got that correct, that would be clean indeed.
>>>>
>>>> But then we would have a different encoding for block entries and page
>>>> entries. In page entries, DBM could be used to say if the page is
>>>> writable,
>>>> but on block entries one would have to look at the 'dirty-bit'.
>>>>
>>>> Would that be ok?
>>>>
>>>> Thanks!
>>>> Leo
>>>>
>>> Hi Leo,
>>>
>>> My initial concern was that clearing all DBM bits at the start of
>>> migration would be too expensive, so I thought distinguishing between
>>> level-3 entries and block entries would be better.
>>>
> 
> I think we expect it to be expensive, but since we already clean the
> dirty-bit (ro/rw) bit, we can have both happening in the same write :)
> 
> (since we only mark the DBM bit when we fault the memory on lazy-splitting,
> we are expecting to have the same amount of writes to pagetable as we have
> before HDBSS, both on faulting and 1st iteration cleaning)
> 
Hi, Leo

Actually, I have thought about this approach too, but if we clear DBM in
kvm_pgtable_stage2_wrprotect(), then during the first round of
migration, we will fault and release RO -> W, and then add DBM.

But next time, when we migrate the dirty pages in round two, we will run
kvm_pgtable_stage2_wrprotect() again, which will clear DBM again. And
finally, HDBSS will be useless during migration.

> 
>>> However, I ran a quick test on a 400GB VM (4 vCPUs), and the overhead
>>> turned out to be around 30ns — which I think is acceptable.
>>
>> Just a quick correction — I misstated the unit in my previous email. The
>> overhead for clearing DBM on the 400GB VM (4 vCPUs) was around 32 µs, not 30
>> ns.
>>
> 
> Oh, that seems more likely :)
> 
> Question: is tha above amount of memory initially in Level-1 blocks,
> level-2 blocks or level-3 pages? (aka: were you using explicit/transparent
> hugepages?)
> 

I'm using transparent hugepages. However, if we were to use level-3 
stage-2 pages with -mem-prealloc enabled in QEMU, I believe the time 
cost would be extremely high — potentially out of our control.

>>>
>>> So I think we can go with your approach: simply clear DBM globally in
>>> kvm_arch_commit_memory_region() when dirty logging starts, before write-
>>> protecting the memslot.
>>>
>>> ```
>>> void kvm_arch_commit_memory_region(...)
>>> {
>>>       // ...
>>>       if (log_dirty_pages) {
>>>           if (change == KVM_MR_DELETE)
>>>               return;
>>>
>>>           kvm_mmu_clear_dbm_memory_region(kvm, new->id);
> 
> Agree, but see the comment above about using the same write that already
> exists, then we are not supposed to see much of a change.

Please see the discussion above.

> 
> Thanks!
> Leo
> 

Thanks!
Tian