Re: [syzbot] [mm?] INFO: rcu detected stall in kvm_vm_release (4)

Ye Liu <[email protected]>
Newsgroups org.kvack.linux-mm,org.kernel.vger.linux-kernel
Message-ID <[email protected]>

在 2026/8/21 00:29, Andrew Morton 写道:
> On Thu, 20 Aug 2026 06:16:25 -0700 syzbot <[email protected]> wrote:
> 
>> Hello,
>>
>> syzbot found the following issue on:
>>
>> HEAD commit:    e8bf40d15402 Merge tag 'chrome-platform-firmware-v7.3' of ..
>> git tree:       upstream
>> console output: https://syzkaller.appspot.com/x/log.txt?x=1434f815580000
>> kernel config:  https://syzkaller.appspot.com/x/.config?x=1941312e3e971b07
>> dashboard link: https://syzkaller.appspot.com/bug?extid=5f0b6e2e4195c75b5658
>> compiler:       Debian clang version 22.1.8 (++20260613092233+e80beda6e255-1~exp1~20260613092250.77), Debian LLD 22.1.8
>>
>> Unfortunately, I don't have any reproducer for this issue yet.
>>
>> Downloadable assets:
>> disk image: https://storage.googleapis.com/syzbot-assets/7df7b958efe0/disk-e8bf40d1.raw.xz
>> vmlinux: https://storage.googleapis.com/syzbot-assets/cefbf90e524a/vmlinux-e8bf40d1.xz
>> kernel image: https://storage.googleapis.com/syzbot-assets/37a530b91001/bzImage-e8bf40d1.xz
>>
>> IMPORTANT: if you fix the issue, please add the following tag to the commit:
>> Reported-by: [email protected]
>>
>> rcu: INFO: rcu_preempt detected stalls on CPUs/tasks:
>> rcu: 	Tasks blocked on level-0 rcu_node (CPUs 0-1): P21500/1:b..l
>> rcu: 	(detected by 0, t=10503 jiffies, g=207513, q=1145 ncpus=2)
>> task:syz.9.4605      state:R  running task     stack:26920 pid:21500 tgid:21500 ppid:14397  task_flags:0x40044c flags:0x00080003
> 
> Thanks.
> 
> lazybutt pasted this into Gemini:
> 	https://share.gemini.google/9rAubqC4qNzf
> 
> : The combination of debug tools: KASAN quarantine processing + ORC
> : unwinder + page_owner stack tracing executed so slowly on the CPU that
> : it exceeded the RCU stall threshold.
> 
> (I think we've recently seen the same thing, different callers)
> 
> I can't say I'm terribly motivated.  page_owner is expensive and if you
> run heavy stress testing with page_owner enabled then such things
> aren't surprising.
> 
> Let me Cc a few page_owner and KASAN people, see if they feel
> differently.

Adding cond_resched() to the qlist_free_all() loop would fix this .
kasan_quarantine_remove_cache() already dose the same in its scanning loop.
> 
>> Call Trace:
>>  <TASK>
>>  context_switch kernel/sched/core.c:5510 [inline]
>>  __schedule+0x17d4/0x5630 kernel/sched/core.c:7239
>>  preempt_schedule_irq+0x4b/0x90 kernel/sched/core.c:7561
>>  irqentry_exit_to_kernel_mode include/linux/irq-entry-common.h:539 [inline]
>>  irqentry_exit+0x14f/0x910 kernel/entry/common.c:167
>>  asm_sysvec_apic_timer_interrupt+0x1a/0x20 arch/x86/include/asm/idtentry.h:674
>> RIP: 0010:lock_acquire+0x222/0x350 kernel/locking/lockdep.c:5890
>> Code: ff ff ff e8 c0 0b 32 0a f7 44 24 08 00 02 00 00 0f 84 39 ff ff ff 65 48 8b 05 12 77 cb 11 48 3b 44 24 58 75 33 fb 48 83 c4 60 <5b> 41 5c 41 5d 41 5e 41 5f 5d c3 cc cc cc cc cc 48 8d 3d 57 a1 b8
>> RSP: 0018:ffffc900074cf260 EFLAGS: 00000286
>> RAX: 7db047bb98fc4c00 RBX: 0000000000000000 RCX: 0000000000000046
>> RDX: 00000000c0065176 RSI: ffffffff8e4b5735 RDI: ffffffff8c4bdd80
>> RBP: ffffffff8177710f R08: ffffffff8177710f R09: 0000000000000000
>> R10: 0000000000000000 R11: ffffffff8eb5a0e0 R12: 0000000000000002
>> R13: ffffffff8eb5a0e0 R14: 0000000000000000 R15: 0000000000000246
>>  rcu_lock_acquire include/linux/rcupdate.h:300 [inline]
>>  rcu_read_lock include/linux/rcupdate.h:840 [inline]
>>  class_rcu_constructor include/linux/rcupdate.h:1183 [inline]
>>  unwind_next_frame+0xac/0x2550 arch/x86/kernel/unwind_orc.c:495
>>  arch_stack_walk+0x11b/0x150 arch/x86/kernel/stacktrace.c:25
>>  stack_trace_save+0xa9/0x100 kernel/stacktrace.c:122
>>  save_stack+0x122/0x230 mm/page_owner.c:165
>>  __reset_page_owner+0x71/0x1f0 mm/page_owner.c:320
>>  reset_page_owner include/linux/page_owner.h:25 [inline]
>>  __free_pages_prepare mm/page_alloc.c:1406 [inline]
>>  __free_frozen_pages+0xc22/0xd20 mm/page_alloc.c:2950
>>  __slab_free+0x274/0x2c0 mm/slub.c:5741
>>  qlink_free mm/kasan/quarantine.c:163 [inline]
>>  qlist_free_all+0x99/0x100 mm/kasan/quarantine.c:179
>>  kasan_quarantine_reduce+0x148/0x160 mm/kasan/quarantine.c:286
>>  __kasan_slab_alloc+0x22/0x80 mm/kasan/common.c:350
>>  kasan_slab_alloc include/linux/kasan.h:253 [inline]
>>  slab_post_alloc_hook mm/slub.c:4584 [inline]
>>  slab_alloc_node mm/slub.c:4917 [inline]
>>  __do_kmalloc_node mm/slub.c:5333 [inline]
>>  __kmalloc_noprof+0x315/0x720 mm/slub.c:5359
>>  _kmalloc_noprof include/linux/slab.h:992 [inline]
>>  _kzalloc_noprof include/linux/slab.h:1309 [inline]
>>  kobject_get_path+0xc5/0x2f0 lib/kobject.c:161
>>  kobject_uevent_env+0x29e/0x9e0 lib/kobject_uevent.c:548
>>  kvm_uevent_notify_change+0x300/0x3a0 virt/kvm/kvm_main.c:6335
>>  kvm_destroy_vm virt/kvm/kvm_main.c:1260 [inline]
>>  kvm_put_kvm+0xa9/0xb50 virt/kvm/kvm_main.c:1331
>>  kvm_vm_release+0x43/0x50 virt/kvm/kvm_main.c:1354
>>  __fput+0x418/0xa50 fs/file_table.c:512
>>  task_work_run+0x1d9/0x270 kernel/task_work.c:233
>>  exit_task_work include/linux/task_work.h:40 [inline]
>>  do_exit+0x73a/0x2360 kernel/exit.c:1009
>>  do_group_exit+0x22d/0x2f0 kernel/exit.c:1152
>>  get_signal+0x121b/0x12c0 kernel/signal.c:3046
>>  arch_do_signal_or_restart+0xbb/0x860 arch/x86/kernel/signal.c:337
>>  __exit_to_user_mode_loop kernel/entry/common.c:66 [inline]
>>  exit_to_user_mode_loop+0x10e/0x770 kernel/entry/common.c:101
>>  __exit_to_user_mode_prepare include/linux/irq-entry-common.h:207 [inline]
>>  syscall_exit_to_user_mode_prepare include/linux/irq-entry-common.h:230 [inline]
>>  syscall_exit_to_user_mode include/linux/entry-common.h:336 [inline]
>>  do_syscall_64+0x328/0x520 arch/x86/entry/syscall_64.c:89
>>  entry_SYSCALL_64_after_hwframe+0x77/0x7f
>> RIP: 0033:0x7fe3a1f5e90e
>> RSP: 002b:00007fe3a2e23f58 EFLAGS: 00000246 ORIG_RAX: 00000000000000e6
>> RAX: fffffffffffffdfc RBX: 00007fe3a2e246c0 RCX: 00007fe3a1f5e90e
>> RDX: 00007fe3a2e23fb0 RSI: 0000000000000000 RDI: 0000000000000000
>> RBP: 00007fe3a2035024 R08: 0000000000000000 R09: 0000000000000000
>> R10: 0000000000000000 R11: 0000000000000246 R12: 0000000000000000
>> R13: 00007fe3a2226308 R14: 00007fe3a2226270 R15: 00007fe3a234fa48
>>  </TASK>
>>
>>
>> ---
>> This report is generated by a bot. It may contain errors.
>> See https://goo.gl/tpsmEJ for more information about syzbot.
>> syzbot engineers can be reached at [email protected].
>>
>> syzbot will keep track of this issue. See:
>> https://goo.gl/tpsmEJ#status for how to communicate with syzbot.
>>
>> If the report is already addressed, let syzbot know by replying with:
>> #syz fix: exact-commit-title
>>
>> If you want to overwrite report's subsystems, reply with:
>> #syz set subsystems: new-subsystem
>> (See the list of subsystem names on the web dashboard)
>>
>> If the report is a duplicate of another one, reply with:
>> #syz dup: exact-subject-of-another-report
>>
>> If you want to undo deduplication, reply with:
>> #syz undup

-- 
Thanks,
Ye Liu
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.