Re: [PATCH 0/5] x86/mm/pat: CPA fixes
Pedro Falcato <[email protected]> Mon, 3 Aug 2026 13:41:37 +0100
| Newsgroups | dev.linux.lists.iommu,org.kernel.vger.linux-kernel,org.kernel.vger.stable,org.kvack.linux-mm |
|---|---|
| Message-ID | <anCK3eWFMwZqq5ka@pedro-suse> |
--bh5jbue2ogdsvams Content-Type: text/plain; charset=us-ascii Content-Disposition: inline On Thu, Jul 30, 2026 at 05:53:50PM +0200, Steffen Dirkwinkel wrote: > Hello, > > > On Tue, 2026-07-28 at 16:07 +0300, Mike Rapoport (Microsoft) wrote: > > There are a couple of CPA fixes floating around: > > > > Denis Lunev fixed races between split and collapse of the large mappings: > > > > https://lore.kernel.org/all/[email protected] > > We saw the error below and I was wondering if it might be related to these fixes > or a similar case that's unfixed still. Seems to have happened during concurrent > kernel module loading of kvm and i915 (similar to the case in the patch from > Denis Lunev). But we got it without KASAN and the stack looks a little > different. I was not able to reproduce this with a ~16 hour concurrent module > load unload loop so far. > > Kernel: v7.1.5, PREEMPT_RT, tainted because of /dev/msr access > CPU: Elkhart Lake Atom X6214RE, 2 cores, isolcpus=1-N > > ------------[ cut here ]------------ > kernel BUG at arch/x86/kernel/alternative.c:2644! > Oops: invalid opcode: 0000 [#1] SMP NOPTI > CPU: 0 UID: 0 PID: 559 Comm: (udev-worker) Tainted: G S 7.1.5- > rt1-bhf-369933-f1a4ee1dd787 #1 PREEMPT_{RT,(lazy)} > Tainted: [S]=CPU_OUT_OF_SPEC > RIP: 0010:__text_poke+0x356/0x3d0 > Call Trace: > <TASK> > smp_text_poke_batch_finish+0x1aa/0x3a0 > ? vmx_switch_vmcs+0xc8/0xd0 [kvm_intel] > __static_call_transform+0xfa/0x1f0 > ? vmx_switch_vmcs+0xc8/0xd0 [kvm_intel] > ? __pfx_preempt_schedule_thunk+0x10/0x10 > arch_static_call_transform+0x57/0xa0 > ? vmx_switch_vmcs+0xc8/0xd0 [kvm_intel] > __static_call_init+0x1aa/0x230 > ? __SCT__tp_func_kvm_mmu_split_huge_page+0x8/0x8 > ? __SCT__tp_func_kvm_mmu_split_huge_page+0x8/0x8 > static_call_module_notify+0x11f/0x150 We're also observing something very similar downstream (https://bugzilla.opensuse.org/show_bug.cgi?id=1271202) on 7.1.3 kernels, on our QA infra. I wrote a patch that should fix it (I didn't get to test it yet, since this didn't hit mainline/stable yet), on top of this. Mike, do you mind throwing it on top of this series, or should I submit separately? -- Pedro --bh5jbue2ogdsvams Content-Type: text/x-patch; charset=us-ascii Content-Disposition: attachment; filename="0001-x86-alternative-exclude-text-poking-against-change_p.patch" From 0a483fd05fda6e3953861b722995b58daf4878ef Mon Sep 17 00:00:00 2001 From: Pedro Falcato <[email protected]> Date: Mon, 3 Aug 2026 13:16:25 +0100 Subject: [PATCH] x86/alternative: exclude text poking against change_page_attr() From time to time, the following BUG can be observed[0]: > kernel BUG at arch/x86/kernel/alternative.c:2576! > Oops: invalid opcode: 0000 [#1] SMP NOPTI > CPU: 0 UID: 0 PID: 355 Comm: (udev-worker) Not tainted 7.1.3-1-default #1 PREEMPT(full) openSUSE Tumbleweed 8c1795b03ec64f997e57a8ad38b1161e3b98da64 > Hardware name: QEMU Standard PC (i440FX + PIIX, 1996), BIOS unknown 02/02/2022 > RIP: 0010:__text_poke+0x2aa/0x450 > Call Trace: > <TASK> > smp_text_poke_batch_finish+0x2a7/0x320 > __static_call_transform+0xb7/0x220 > arch_static_call_transform+0x5b/0xb0 > __static_call_init+0xe9/0x270 > static_call_module_notify+0x11f/0x150 > notifier_call_chain+0x61/0xe0 > blocking_notifier_call_chain_robust+0x63/0xc0 > load_module+0x1c92/0x20c0 > init_module_from_file+0xd8/0x140 > idempotent_init_module+0x100/0x2f0 > __x64_sys_finit_module+0x71/0xe0 > do_syscall_64+0xe1/0x610 > entry_SYSCALL_64_after_hwframe+0x76/0x7e which matches the following BUG_ON in alternative.c: /* * If something went wrong, crash and burn since recovery paths are not * implemented. */ BUG_ON(!pages[0] || (cross_page_boundary && !pages[1])); This can happen if vmalloc_to_page() fails, for any reason. Such can happen if text poking races with CPA, which can possibly result in the collapsing of page tables (or breaking of PMD hugepages). It is not a problem for most users of vmalloc_to_page() (they solely own the vmalloc'd range) but, when CONFIG_ARCH_HAS_EXECMEM_ROX=y, various modules own a single execmem vmalloc range, and can call set_memory_*() in parallel on it. This can happen to race against __text_poke and cause havoc in vmalloc_to_page(). Fix it by excluding against CPA using the init_mm mmap read lock. Fixes: 64f6a4e10c05 ("x86: re-enable EXECMEM_ROX support") Reported-by: Jiri Slaby <[email protected]> Link: https://bugzilla.opensuse.org/show_bug.cgi?id=1271202 [0] Reported-by: Steffen Dirkwinkel <[email protected]> Link: https://lore.kernel.org/linux-mm/[email protected]/ Cc: [email protected] Signed-off-by: Pedro Falcato <[email protected]> --- arch/x86/kernel/alternative.c | 8 ++++++++ 1 file changed, 8 insertions(+) diff --git a/arch/x86/kernel/alternative.c b/arch/x86/kernel/alternative.c index 62936a3bde19..9071eb870eab 100644 --- a/arch/x86/kernel/alternative.c +++ b/arch/x86/kernel/alternative.c @@ -2559,6 +2559,14 @@ static void *__text_poke(text_poke_f func, void *addr, const void *src, size_t l */ BUG_ON(!after_bootmem); + /* + * Exclude against change_page_attr() collapse in execmem ROX regions. + * These are PMD sized and this module may not own the whole PMD, + * thus breakdown/collapse may happen at any moment by concurrent module + * loading, which races with vmalloc_to_page(). + */ + guard(mmap_read_lock)(&init_mm); + if (!core_kernel_text((unsigned long)addr)) { pages[0] = vmalloc_to_page(addr); if (cross_page_boundary) -- 2.55.0 --bh5jbue2ogdsvams--