Re: [PATCH 1/2] mm/damon/core: cover discrete System RAM areas with per-range regions
Jiayuan Chen <[email protected]> Tue, 28 Jul 2026 23:17:01 +0800
| Newsgroups | dev.linux.lists.damon,org.kernel.vger.linux-doc,org.kernel.vger.linux-kernel,org.kvack.linux-mm |
|---|---|
| Message-ID | <[email protected]> |
On 7/28/26 10:30 PM, SJ Park wrote: > On Tue, 28 Jul 2026 18:06:48 +0800 Jiayuan Chen <[email protected]> wrote: > >> On 7/27/26 10:26 PM, SJ Park wrote: >>> Hello Jiayuan, >>> >>> On Mon, 27 Jul 2026 17:54:22 +0800 Jiayuan Chen <[email protected]> wrote: >>> >>>> From: Jiayuan Chen <[email protected]> >>>> >>>> damon_set_region_system_rams_default(), introduced by commit 70d8797c15d6 >>>> ("mm/damon: introduce damon_set_region_system_rams_default()"), is used by >>>> DAMON_RECLAIM, DAMON_LRU_SORT and DAMON_STAT to set the default monitoring >>>> target address range covering all 'System RAM' when the user does not >>>> specify a range. It walks the 'System RAM' resources but keeps only the >>>> start of the first resource and the end of the last one, and then sets a >>>> single monitoring region spanning that whole [first_start, last_end] range. >>>> >>>> On systems whose RAM is split into discrete areas that are far apart in the >>>> physical address space, that single region also covers the holes between >>>> them. For example: >>>> >>>> $ sudo cat /proc/iomem | grep RAM >>>> 00001000-0009ffff : System RAM >>>> 00100000-4848c017 : System RAM >>>> 4848c018-48550c57 : System RAM >>>> 48550c58-48551017 : System RAM >>>> 48551018-48615c57 : System RAM >>>> 48615c58-48616017 : System RAM >>>> 48616018-486dac57 : System RAM >>>> 486dac58-4e563017 : System RAM >>>> 4e563018-4e627c57 : System RAM >>>> 4e627c58-4ef39017 : System RAM >>>> 4ef39018-4ef3f057 : System RAM >>>> 4ef3f058-4efe6017 : System RAM >>>> 4efe6018-4efec057 : System RAM >>>> 4efec058-50247fff : System RAM >>>> 50317000-56720fff : System RAM >>>> 56722000-59c19fff : System RAM >>>> 6bbfe000-6bbfefff : System RAM >>>> 6bc00000-777fffff : System RAM >>>> 100000000-1007effffff : System RAM >>>> 67e80000000-77e7fffffff : System RAM >>>> >>>> Here the last two areas (about 1TB starting at 4GiB, and about 1.1TB >>>> starting at ~6.5TB) are separated by a ~5.5TB hole, and the single-region >>>> setup makes DAMON treat that entire hole as if it were memory. >>>> >>>> This is harmful in a few ways. The monitoring target regions are limited >>>> by max_nr_regions, so regions that fall into the hole waste that budget and >>>> leave fewer regions for the real RAM, coarsening the adaptive regions and >>>> degrading the monitoring accuracy. >>> The hole would look like not accessed. As a result, the whole region will be a >>> few regions that very cold. That wouldn't waste the budget that much. Do you >>> have some specific setups that this cannot help? >> It's my mistake. >> >> I first saw the kdamond CPU go up, and I assumed it was the number of >> regions. I was wrong. >> >> I re-tested and profiled it with perf. It's not the region count — it's >> the page-by-page walk of the hole. >> >> On a VM with a 116GiB hole and a single [first,last] span, I enabled >> DAMON_LRU_SORT and ran perf on its kdamond: >> 30.15% damon_get_folio >> 17.09% pfn_to_online_page >> 2.24% __nr_to_section >> ... >> 0.07% damon_split_region_at >> 0.06% damon_merge_two_regions >> >> Almost all of it is the per-page folio lookup. Merge and split are ~0.1%. >> The cold scheme targets cold, old regions. The hole is never accessed, so >> it's always cold and old, and it matches. Then the action walks its whole >> empty span page by page, every apply. RECLAIM and LRU_SORT do this by >> default. It scales with the hole size, so on the real 5.5TiB machine the >> kdamond can't keep up. > Makes sense, thank you for investigating and sharing this! > >> >>>> In addition, DAMOS actions on the paddr >>>> operations set walk such a region page by page, so a region that covers the >>>> hole is walked for its entire (empty) span on every application. >>> That makes sense. >>> >>>> Set a separate monitoring region for each discrete System RAM area instead, >>>> coalescing only truly adjacent (no gap in between) resources into one >>>> range, so holes between the areas are excluded. The reported *start and >>>> *end still carry the overall first-start and last-end, so the user-visible >>>> default range reported via the module parameters is unchanged. >>> I'm concerned if this could result in having too many regions. The gap between >> On the region count: can't we just coalesce any hole smaller than >> min_region_sz? > I'm not fully understanding your point. min_region_sz is only 4 KiB by > default, and anyway DAMON cannot create regions of size smaller than > min_region_sz. I cannot get how this helps. > >> After that, the number of ranges is just the number of >> discrete >> System RAM areas > I'm still not convinced with this. What if the number of discrete system ram > areas is larger than max_nr_regions? > > Actually I was also thinking about this problem for in the past. One of my > idea at that time was, handle only a few largest holes that practically being > problems. That is, while reading the system ram layout, find the holes, sort > those by size, and do make holes in DAMON regions layout for the biggest N > (say, 2) holes. This may handle most cases including your 5 TiB hole. > Actually vaddr is doing this, so we may be able to reuse some of the code. > > If it makes sense to you, I will try to implement this. One worry with a fixed N is that it's a magic number. If a machine has more than N big holes (more sockets / NUMA nodes, or several CXL devices), the extra ones are silently not excluded and get walked again. And making N configurable just moves the per-machine tuning back to the user, which is exactly what I'm trying to avoid. vaddr over-includes the gaps into its regions too, but it can skip them cheaply: it walks the VMA tree (find_vma / maple tree), it's cheap. paddr has no such structure, so an included hole is walked in full. That's why a fixed N is safe for vaddr but risky for paddr. >>> user-visible parameters and internal state is also a concern. >> >> I think monitor_region_start/end is meant to expose the overall range, and >> skipping the holes inside it is an implementation detail. Even if we didn't >> skip them, the adaptive region count is never 1 anyway — DAMON already >> splits [first,last] into many regions. The holes just add a few more, so >> the param never matched the internal state exactly to begin with. > It is arguable, but I believe this also makes sense in my opinion. > >> >>> A quick workaround would be adjusting the memory layout in BIOS, using DAMON >>> sysfs interface instead, or setting the monitor_region_{start,end} to cover >>> only the single area. Have you considered such workarounds? >>> >>> Let's complete this high level discussion first. >> >> Right, this is doable today — the DAMON sysfs interface can set multiple >> >> regions manually. This patch is only about convenience. > Are you actually running DAMON_LRU_SORT or DAMON_RECLAIM in a production system > having 5 TiB hole? Or, planning to do? If the above idea makes sense to you, > I could prioritize implementation of it depending on this. I'm not actually running DAMON_LRU_SORT or DAMON_RECLAIM on such a machine. I ran into this while working on per-cgroup hot/cold page tracking, where I was comparing performance and accuracy across setups — that's where these numbers came from. > > Thanks, > SJ > > [...] And to be clear, I'm not trying to land a patch here — I just noticed this as a possible optimization. If you have a better idea, I'm happy to go with it :).