Re: [PATCH 03/22] mm: introduce MMF_KERNEL flag and set it for init_mm

David Laight <[email protected]>
Newsgroups org.kernel.vger.linux-parisc,dev.linux.lists.loongarch,org.infradead.lists.linux-riscv,org.infradead.lists.linux-um,org.kernel.vger.linux-arch,org.kernel.vger.linux-efi,org.kernel.vger.linux-kernel,org.kernel.vger.linux-s390,org.kernel.vger.sparclinux,org.kvack.linux-mm,org.ozlabs.lists.linuxppc-dev
Message-ID <20260811121941.4ca20b88@pumpkin>
On Tue, 11 Aug 2026 11:00:09 +0200
Kevin Brodsky <[email protected]> wrote:

> On 11/08/2026 10:10, David Laight wrote:
> > On Wed, 5 Aug 2026 14:07:01 +0200
> > "David Hildenbrand (Arm)" <[email protected]> wrote:
> >  
> >> On 8/3/26 15:59, Christophe Leroy (CS GROUP) wrote:  
> >>>
> >>> Le 14/07/2026 à 16:03, Kevin Brodsky a écrit :    
> >>>> mm code often needs to know whether some mm represents a kernel or
> >>>> user address space. This is currently done by comparing the mm
> >>>> pointer with &init_mm; besides not being particularly elegant, this
> >>>> ignores the fact that other mm's (e.g. efi_mm) may also represent
> >>>> parts of the kernel address space.
> >>>>
> >>>> Introduce a new mm flag MMF_KERNEL and set it for init_mm.
> >>>> Subsequent patches will use this flag to replace comparisons with
> >>>> &init_mm. No functional change is introduced for now.    
> >>> Did you consider performance impact ? This test is usually done in quite
> >>> critical memory handling functions.
> >>>
> >>> init_mm is known at link time. Before your patch 08/22 there is just a
> >>> comparison of mm (r3) with a constant (loaded in r10):
> >>>
> >>> c0014048 <assert_pte_locked>:
> >>> c0014048:    3d 40 c1 09     lis     r10,-16119
> >>> c001404c:    39 4a 03 98     addi    r10,r10,920
> >>> c0014050:    7c 03 50 00     cmpw    r3,r10
> >>> c0014054:    4d 82 00 20     beqlr
> >>> ...
> >>>
> >>> After patch 08/22 we have, it first checks that mm is not 0, then it loads the
> >>> word located at mm+528 then AND it with 0x1. This load might be costly.
> >>>
> >>> c0014048 <assert_pte_locked>:
> >>> c0014048:    2c 03 00 00     cmpwi   r3,0
> >>> c001404c:    7c 85 23 78     mr      r5,r4
> >>> c0014050:    41 82 00 10     beq     c0014060 <assert_pte_locked+0x18>
> >>> c0014054:    81 23 02 10     lwz     r9,528(r3)
> >>> c0014058:    71 29 00 01     andi.   r9,r9,1
> >>> c001405c:    4c 82 00 20     bnelr    
> >> Is that a real problem, though?  
> > The cost of the read is likely to matter most if the branch gets mispredicted.
> > I'd also guess the mm isn't usually NULL - so that branch needs to statically
> > predicted correctly as well.
> > I'd also not assume that the mm is in the cache (unless the surrounding code
> > has already accessed it), the L1 data caches are small.  
> 
> You're right, I didn't think about this the right way round - in general
> it may not be easy to predict that mm points to init_mm, and an
> arbitrary user mm isn't guaranteed to be cached. Repeated calls to
> mm_is_kernel(mm) in the same function shouldn't cost much though.

Depends.
If it is real function call the code will be in the I-cache and is likely
to get predicted correctly by the third call.
OTOH if inlined you've all the I-cache loads and that static branch
prediction might get them all wrong.

Oh, and saving the result in a local might cause a register spill.

	David

> 
> Would some (micro)benchmark be particularly relevant to measure the
> overhead, if any?
> 
> - Kevin
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.