[PATCH v4] sched/fair: Preserve wake-affine CPU for non-SMT reciprocal sync wakeups

"Shubhang Kaushik (Ampere)" <[email protected]> Mon, 03 Aug 2026 18:45:49 -0700
Newsgroups gmane.linux.kernel
Message-ID <[email protected]>
For WF_SYNC wakeups, wake_affine() may select the waker CPU, but the CFS
wakeup path still passes that target to select_idle_sibling(). The idle
CPU search can then move the wakee away from the wake-affine target.

Pipe-style ping-pong workloads expose this because the wakee is handed
back and forth between two tasks. In that case, moving the wakee to
another idle CPU can cost more than preserving the wake-affine waker CPU.

Use the existing last_wakee and wake_wide() state to identify narrow
reciprocal WF_SYNC wakeups:

  A wakes B
  B wakes A
  A wakes B
  ...

Handle only this narrow reciprocal case on non-SMT systems. Once the
wake-affine path has selected or kept the waker CPU, preserve that target
when the waker rq has no other runnable fair task. Return the waker CPU
before select_idle_sibling() so the idle CPU search does not move this
handoff away from the wake-affine target.

This does not define a generic WF_SYNC placement rule. Generic WF_SYNC
wakeups continue through the existing wake_affine() and
select_idle_sibling() behavior. SMT systems also continue through
select_idle_sibling(), where idle sibling/core placement can be handled
with SMT topology visible.

On asymmetric-capacity systems, still require the wakee to fit on the
waker CPU.

Signed-off-by: Shubhang Kaushik (Ampere) <[email protected]>
---
Tested on 80-core non-SMT Ampere Altra, tip:sched/core baseline.

perf bench sched pipe -l 1000000, 20 runs:
default:
  3.985 -> 3.187 usec/op mean,   about 20.0% improvement
  4.026 -> 3.181 usec/op median, about 21.0% improvement

taskset -c 78,79:
  3.851 -> 3.144 usec/op mean,   about 18.4% improvement
  3.804 -> 3.140 usec/op median, about 17.4% improvement

taskset -c 79:
  3.055 -> 3.113 usec/op mean,   about 1.9% slower
  3.045 -> 3.109 usec/op median, about 2.1% slower

Hackbench process/thread pipe cases with 1/2/4/8 groups were within
noise, with mean deltas from -1.8% to +3.7% over 10 runs.

Schbench normal mode at 8/40/80/240 workers and schbench pipe mode at
1/2/4/8 workers showed no material regression.

Baseline: tip/sched/core at 5186ef36909c
---
Changes in v4:
  - Preserve the waker CPU only after the wake-affine path selected or
    kept it.
  - Clarify that WF_SYNC remains a hint, not a generic placement rule.
  - Leave SMT systems on the existing select_idle_sibling() path.
  - Refresh testing on tip:sched/core.

Link to v3: https://lore.kernel.org/r/[email protected]

Changes in v3:
  - Limit the direct waker-CPU preference to !sched_smt_active(); SMT
    systems continue through the existing wake_affine() and
    select_idle_sibling() path.
  - Drop the redundant affinity check; want_affine already verifies the
    waker CPU is allowed.
  - Use a plain p->last_wakee read instead of READ_ONCE().
  - Rebase and refresh testing on v7.2-rc5.

Link to v2: https://lore.kernel.org/r/[email protected]

Changes in v2:
  - Move the reciprocal handoff preference under the existing
    SD_WAKE_AFFINE domain check.
  - Drop futex from the changelog motivation.
  - Refresh perf bench sched pipe results after rebasing.

Link to v1: https://lore.kernel.org/r/[email protected]
---
 kernel/sched/fair.c | 26 ++++++++++++++++++++++++++
 1 file changed, 26 insertions(+)

diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index d78467ec6ee1343050fcc2794dafb38ade3599e5..e377dc43460a691315d895e5cde5208987791ce8 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -8794,6 +8794,26 @@ static inline bool asym_fits_cpu(unsigned long util,
 	return true;
 }
 
+/*
+ * For reciprocal WF_SYNC handoffs, prefer the waker CPU when it has no
+ * other runnable fair task.
+ */
+static bool prefer_sync_pair_cpu(struct task_struct *p, int cpu)
+{
+	struct rq *rq = cpu_rq(cpu);
+
+	if ((rq->nr_running - cfs_h_nr_delayed(rq)) != 1)
+		return false;
+
+	if (sched_asym_cpucap_active()) {
+		sync_entity_load_avg(&p->se);
+		if (!task_fits_cpu(p, cpu))
+			return false;
+	}
+
+	return true;
+}
+
 /*
  * Try and locate an idle core/thread in the LLC cache domain.
  */
@@ -9582,6 +9602,12 @@ select_task_rq_fair(struct task_struct *p, int prev_cpu, int wake_flags)
 			if (cpu != prev_cpu)
 				new_cpu = wake_affine(tmp, p, cpu, prev_cpu, sync);
 
+			if (sync && !sched_smt_active() &&
+			    new_cpu == cpu &&
+			    p->last_wakee == current &&
+			    prefer_sync_pair_cpu(p, cpu))
+				return cpu;
+
 			sd = NULL; /* Prefer wake_affine over balance flags */
 			break;
 		}

---
base-commit: 075b74841bd0065a3bda3440873c747938e69b68
change-id: 20260721-b4-sched-sync-wakeup-04d40cbeb1da

Best regards,
-- 
Shubhang Kaushik (Ampere) <[email protected]>