[RFC PATCH 2/7] mm: dispatch __mmput() to an mm_reaper kthread via llist

Aditya Sharma <[email protected]>
Newsgroups org.kernel.vger.linux-trace-kernel,org.kernel.vger.linux-doc,org.kernel.vger.linux-kernel,org.kvack.linux-mm
Message-ID <[email protected]>
Introduce a single dedicated, freezable, sleepable kthread (mm_reaper)
modeled on oom_reaper, and the machinery to route the deferrable unit of
teardown -- __mmput() -- to it through a lock-free llist plus waitqueue.
mmput_exit() does llist_add() and wakes the kthread, which drains with
llist_del_all() and runs __mmput() per entry with cond_resched(). The
drain loop also calls try_to_freeze() after each entry, so a freeze
request is honoured mid-batch rather than only at the wait point.

Nothing calls mmput_exit() yet: exit_mm() is not routed to it, and there
is no eligibility check or gate, so it would queue every final-drop mm
unconditionally if invoked. This changes no behavior -- the function is
built but unreferenced, the kthread receives no work, and __mmput() runs
inline on the exiting CPU exactly as today. The exit_mm() hook and the
eligibility/gating policy follow in later patches.

llist (rather than oom_reaper's spinlock+list) avoids a global lock on
the exit path under many-core parallel exits. llist_for_each_entry_safe()
caches the next pointer before __mmput(), so iteration is safe even when
the final mmdrop() inside __mmput() frees mm.

Lifetime: the mm_count reference that mm_users>0 stands for is released
only by that final mmdrop() inside __mmput(). Deferring __mmput() keeps
the reference held, so the mm_struct and its embedded async_reap_node
stay alive until the kthread runs. No extra mmgrab() is needed, for the
same reason mmput_async() needs none.

The kthread is started at subsys_initcall and set to nice 19 to minimize
competition with other EEVDF tasks. Its affinity is expressed as a
preference via kthread_affine_preferred() over the HK_TYPE_KTHREAD
housekeeping mask, so it keeps off isolcpus= and cpuset-isolated CPUs
without the hard binding kthread_bind_mask() would impose --
PF_NO_SETAFFINITY would otherwise leave admins unable to re-affine it.

Signed-off-by: Aditya Sharma <[email protected]>
---
 kernel/fork.c | 54 +++++++++++++++++++++++++++++++++++++++++++++++++++
 1 file changed, 54 insertions(+)

diff --git a/kernel/fork.c b/kernel/fork.c
index f0e2e131a..884e4e767 100644
--- a/kernel/fork.c
+++ b/kernel/fork.c
@@ -112,6 +112,7 @@
 #include <linux/unwind_deferred.h>
 #include <linux/pgalloc.h>
 #include <linux/uaccess.h>
+#include <linux/sched/isolation.h>
 
 #include <asm/mmu_context.h>
 #include <asm/cacheflush.h>
@@ -3407,3 +3408,56 @@ static int __init init_fork_sysctl(void)
 }
 
 subsys_initcall(init_fork_sysctl);
+
+#ifdef CONFIG_ASYNC_MM_TEARDOWN
+static LLIST_HEAD(mm_reaper_list);
+static DECLARE_WAIT_QUEUE_HEAD(mm_reaper_wait);
+
+static void async_mm_teardown_queue(struct mm_struct *mm)
+{
+	if (llist_add(&mm->async_reap_node, &mm_reaper_list))
+		wake_up(&mm_reaper_wait);
+}
+
+static int mm_reaper(void *unused)
+{
+	set_freezable();
+	while (true) {
+		struct llist_node *batch;
+		struct mm_struct *mm, *n;
+
+		wait_event_freezable(mm_reaper_wait, !llist_empty(&mm_reaper_list));
+		batch = llist_del_all(&mm_reaper_list);
+		llist_for_each_entry_safe(mm, n, batch, async_reap_node) {
+			__mmput(mm); /* may free mm via mmdrop */
+			cond_resched();
+			try_to_freeze();
+		}
+	}
+	return 0;
+}
+
+void mmput_exit(struct mm_struct *mm)
+{
+	might_sleep();
+	if (!atomic_dec_and_test(&mm->mm_users))
+		return;
+	async_mm_teardown_queue(mm);
+}
+
+static int __init mm_reaper_init(void)
+{
+	struct task_struct *th;
+
+	th = kthread_create(mm_reaper, NULL, "mm_reaper");
+	if (IS_ERR(th)) {
+		pr_err("mm_reaper: failed to start kthread: %ld\n", PTR_ERR(th));
+		return PTR_ERR(th);
+	}
+	set_user_nice(th, 19); /* minimize competition with other fair class tasks */
+	kthread_affine_preferred(th, housekeeping_cpumask(HK_TYPE_KTHREAD));
+	wake_up_process(th);
+	return 0;
+}
+subsys_initcall(mm_reaper_init);
+#endif
-- 
2.34.1
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.