Re: [PATCH 0/5] x86/mm/pat: CPA fixes

Pedro Falcato <[email protected]> Mon, 3 Aug 2026 13:41:37 +0100
Newsgroups dev.linux.lists.iommu,org.kernel.vger.linux-kernel,org.kernel.vger.stable,org.kvack.linux-mm
Message-ID <anCK3eWFMwZqq5ka@pedro-suse>
--bh5jbue2ogdsvams
Content-Type: text/plain; charset=us-ascii
Content-Disposition: inline

On Thu, Jul 30, 2026 at 05:53:50PM +0200, Steffen Dirkwinkel wrote:
> Hello,
> 
> 
> On Tue, 2026-07-28 at 16:07 +0300, Mike Rapoport (Microsoft) wrote:
> > There are a couple of CPA fixes floating around:
> > 
> > Denis Lunev fixed races between split and collapse of the large mappings:
> > 
> > https://lore.kernel.org/all/[email protected]
> 
> We saw the error below and I was wondering if it might be related to these fixes
> or a similar case that's unfixed still. Seems to have happened during concurrent
> kernel module loading of kvm and i915 (similar to the case in the patch from
> Denis Lunev). But we got it without KASAN and the stack looks a little
> different. I was not able to reproduce this with a ~16 hour concurrent module
> load unload loop so far.
> 
> Kernel: v7.1.5, PREEMPT_RT, tainted because of /dev/msr access
> CPU: Elkhart Lake Atom X6214RE, 2 cores, isolcpus=1-N
> 
> ------------[ cut here ]------------
> kernel BUG at arch/x86/kernel/alternative.c:2644!
> Oops: invalid opcode: 0000 [#1] SMP NOPTI
> CPU: 0 UID: 0 PID: 559 Comm: (udev-worker) Tainted: G S                  7.1.5-
> rt1-bhf-369933-f1a4ee1dd787 #1 PREEMPT_{RT,(lazy)} 
> Tainted: [S]=CPU_OUT_OF_SPEC
> RIP: 0010:__text_poke+0x356/0x3d0
> Call Trace:
>  <TASK>
>  smp_text_poke_batch_finish+0x1aa/0x3a0
>  ? vmx_switch_vmcs+0xc8/0xd0 [kvm_intel]
>  __static_call_transform+0xfa/0x1f0
>  ? vmx_switch_vmcs+0xc8/0xd0 [kvm_intel]
>  ? __pfx_preempt_schedule_thunk+0x10/0x10
>  arch_static_call_transform+0x57/0xa0
>  ? vmx_switch_vmcs+0xc8/0xd0 [kvm_intel]
>  __static_call_init+0x1aa/0x230
>  ? __SCT__tp_func_kvm_mmu_split_huge_page+0x8/0x8
>  ? __SCT__tp_func_kvm_mmu_split_huge_page+0x8/0x8
>  static_call_module_notify+0x11f/0x150

We're also observing something very similar downstream
(https://bugzilla.opensuse.org/show_bug.cgi?id=1271202) on 7.1.3 kernels, on our QA infra.

I wrote a patch that should fix it (I didn't get to test it yet, since
this didn't hit mainline/stable yet), on top of this. Mike, do you mind
throwing it on top of this series, or should I submit separately?

-- 
Pedro

--bh5jbue2ogdsvams
Content-Type: text/x-patch; charset=us-ascii
Content-Disposition: attachment;
	filename="0001-x86-alternative-exclude-text-poking-against-change_p.patch"

From 0a483fd05fda6e3953861b722995b58daf4878ef Mon Sep 17 00:00:00 2001
From: Pedro Falcato <[email protected]>
Date: Mon, 3 Aug 2026 13:16:25 +0100
Subject: [PATCH] x86/alternative: exclude text poking against
 change_page_attr()

From time to time, the following BUG can be observed[0]:

> kernel BUG at arch/x86/kernel/alternative.c:2576!
> Oops: invalid opcode: 0000 [#1] SMP NOPTI
> CPU: 0 UID: 0 PID: 355 Comm: (udev-worker) Not tainted 7.1.3-1-default #1 PREEMPT(full) openSUSE Tumbleweed  8c1795b03ec64f997e57a8ad38b1161e3b98da64
> Hardware name: QEMU Standard PC (i440FX + PIIX, 1996), BIOS unknown 02/02/2022
> RIP: 0010:__text_poke+0x2aa/0x450
> Call Trace:
>  <TASK>
>  smp_text_poke_batch_finish+0x2a7/0x320
>  __static_call_transform+0xb7/0x220
>  arch_static_call_transform+0x5b/0xb0
>  __static_call_init+0xe9/0x270
>  static_call_module_notify+0x11f/0x150
>  notifier_call_chain+0x61/0xe0
>  blocking_notifier_call_chain_robust+0x63/0xc0
>  load_module+0x1c92/0x20c0
>  init_module_from_file+0xd8/0x140
>  idempotent_init_module+0x100/0x2f0
>  __x64_sys_finit_module+0x71/0xe0
>  do_syscall_64+0xe1/0x610
>  entry_SYSCALL_64_after_hwframe+0x76/0x7e

which matches the following BUG_ON in alternative.c:
	/*
	 * If something went wrong, crash and burn since recovery paths are not
	 * implemented.
	 */
	BUG_ON(!pages[0] || (cross_page_boundary && !pages[1]));

This can happen if vmalloc_to_page() fails, for any reason. Such can happen
if text poking races with CPA, which can possibly result in the collapsing
of page tables (or breaking of PMD hugepages). It is not a problem for most
users of vmalloc_to_page() (they solely own the vmalloc'd range) but, when
CONFIG_ARCH_HAS_EXECMEM_ROX=y, various modules own a single execmem vmalloc
range, and can call set_memory_*() in parallel on it. This can happen to
race against __text_poke and cause havoc in vmalloc_to_page().

Fix it by excluding against CPA using the init_mm mmap read lock.

Fixes: 64f6a4e10c05 ("x86: re-enable EXECMEM_ROX support")
Reported-by: Jiri Slaby <[email protected]>
Link: https://bugzilla.opensuse.org/show_bug.cgi?id=1271202 [0]
Reported-by: Steffen Dirkwinkel <[email protected]>
Link: https://lore.kernel.org/linux-mm/[email protected]/
Cc: [email protected]
Signed-off-by: Pedro Falcato <[email protected]>
---
 arch/x86/kernel/alternative.c | 8 ++++++++
 1 file changed, 8 insertions(+)

diff --git a/arch/x86/kernel/alternative.c b/arch/x86/kernel/alternative.c
index 62936a3bde19..9071eb870eab 100644
--- a/arch/x86/kernel/alternative.c
+++ b/arch/x86/kernel/alternative.c
@@ -2559,6 +2559,14 @@ static void *__text_poke(text_poke_f func, void *addr, const void *src, size_t l
 	 */
 	BUG_ON(!after_bootmem);
 
+	/*
+	 * Exclude against change_page_attr() collapse in execmem ROX regions.
+	 * These are PMD sized and this module may not own the whole PMD,
+	 * thus breakdown/collapse may happen at any moment by concurrent module
+	 * loading, which races with vmalloc_to_page().
+	 */
+	guard(mmap_read_lock)(&init_mm);
+
 	if (!core_kernel_text((unsigned long)addr)) {
 		pages[0] = vmalloc_to_page(addr);
 		if (cross_page_boundary)
-- 
2.55.0


--bh5jbue2ogdsvams--