[PATCH 1/1] tcg: defer tb_flush when initial region alloc fails

Yogesh Vyas <[email protected]> Wed, 5 Aug 2026 22:44:53 +0530
Newsgroups gmane.comp.emulators.qemu
Message-ID <[email protected]>
Under MTTCG, busy vCPUs can temporarily exhaust the fixed code-gen
region pool by holding multiple regions. A newly registered vCPU
thread that then fails tcg_region_initial_alloc() currently hits
g_assert and aborts QEMU — seen during vCPU hotplug on ppc64 TCG
(virt-type=qemu) with a large maxcpus.
Instead of asserting, mark the context with tb_flush_pending. The
first tcg_tb_alloc() returns NULL so the existing tb_gen_code path
queues a tb_flush; tcg_region_reset_all() then repacks the pool and
the new thread obtains a region on retry.
This mirrors the graceful exhaustion handling already used by the
runtime tcg_region_alloc() path.

Reported-by: Anushree Mathur <[email protected]>
Closes: https://gitlab.com/qemu-project/qemu/-/work_items/2984
Signed-off-by: Yogesh Vyas <[email protected]>
---
 include/tcg/tcg.h |  9 +++++++++
 tcg/region.c      | 15 ++++++++++++++-
 tcg/tcg.c         | 12 ++++++++++++
 3 files changed, 35 insertions(+), 1 deletion(-)

diff --git a/include/tcg/tcg.h b/include/tcg/tcg.h
index 7669dc1c2d..8a05e3af7b 100644
--- a/include/tcg/tcg.h
+++ b/include/tcg/tcg.h
@@ -381,6 +381,15 @@ struct TCGContext {
     /* Threshold to flush the translated code buffer.  */
     void *code_gen_highwater;
 
+    /*
+     * Set when this context (a newly registered vCPU thread) could not be
+     * given an initial code-gen region because the region pool was
+     * momentarily exhausted by other busy vCPUs. Rather than aborting, the
+     * thread defers a tb_flush: the next tcg_tb_alloc() forces a flush +
+     * retry, which resets the region pool and reclaims slack for this thread.
+     */
+    bool tb_flush_pending;
+
     /* Track which vCPU triggers events */
     CPUState *cpu;                      /* *_trans */
 
diff --git a/tcg/region.c b/tcg/region.c
index 5d4be1453b..1200947385 100644
--- a/tcg/region.c
+++ b/tcg/region.c
@@ -393,7 +393,20 @@ bool tcg_region_alloc(TCGContext *s)
 static void tcg_region_initial_alloc__locked(TCGContext *s)
 {
     bool err = tcg_region_alloc__locked(s);
-    g_assert(!err);
+
+    /*
+     * The region pool can be momentarily exhausted when many busy vCPUs have
+     * each grabbed several regions during translation. A newly started vCPU
+     * thread that cannot get its initial region must not abort QEMU: instead
+     * flag a deferred tb_flush. The thread's first tcg_tb_alloc() will force a
+     * flush + retry (tcg_region_reset_all() resets region.current and
+     * re-packs the live contexts), which reclaims that slack and assigns this
+     * context a region. This mirrors the graceful runtime path in
+     * tcg_tb_alloc()/tcg_region_alloc().
+     */
+    if (err) {
+        s->tb_flush_pending = true;
+    }
 }
 
 void tcg_region_initial_alloc(TCGContext *s)
diff --git a/tcg/tcg.c b/tcg/tcg.c
index 1e77f2365a..ea66fb6f84 100644
--- a/tcg/tcg.c
+++ b/tcg/tcg.c
@@ -1830,6 +1830,18 @@ TranslationBlock *tcg_tb_alloc(TCGContext *s)
     TranslationBlock *tb;
     void *next;
 
+    /*
+     * A newly started vCPU thread that could not obtain its initial region
+     * (region pool momentarily exhausted) deferred a flush. Trigger it now:
+     * returning NULL makes the caller queue a tb_flush and retry, which
+     * resets the region pool and gives this context a valid region. Do this
+     * before touching code_gen_ptr, which is not yet valid for this context.
+     */
+    if (unlikely(s->tb_flush_pending)) {
+        s->tb_flush_pending = false;
+        return NULL;
+    }
+
  retry:
     tb = (void *)ROUND_UP((uintptr_t)s->code_gen_ptr, align);
     next = (void *)ROUND_UP((uintptr_t)(tb + 1), align);
-- 
2.54.0