Re: [PATCH v2 0/3] block: skip the blkcg walk in blk_cgroup_congested() when nothing is throttled
Tejun Heo <[email protected]>
| Newsgroups | org.kernel.vger.linux-block,org.kernel.vger.cgroups,org.kernel.vger.linux-kernel |
|---|---|
| Message-ID | <[email protected]> |
On Fri, Aug 14, 2026 at 09:56:36AM -0700, Usama Arif wrote: > blk_cgroup_congested() walks the current task's blkcg ancestor chain on every > readahead decision and, once swap is in use, on every anonymous and shmem > folio allocation. The answer is almost always "no", but finding that out > costs two loads per level on two cold cache lines, plus an out-of-line > kthread_blkcg() and an RCU read-side pair. On a fleet profile of hosts > running containers with 5-10 level hierarchies it costs about as much as all > of mutex_lock(), 99.4% of it under __folio_throttle_swaprate(). > > Patch 3 gates the walk on a global count of blkcgs with a non-zero > congestion_count, so the common case is a load and a predicted branch. > > That only works if the count is correctly maintained, currently two teardown > paths can leave a blkcg permanently marked congested. Today that only hurts > tasks in the affected cgroup, but it hurts them for the life of the cgroup - > readahead cut to a single page, async readahead skipped, and a throttle > scheduled on every anonymous folio allocation. With a global gate it would > cost every other task on the machine the walk as well. Patches 1 and 2 fix > those two paths and stand on their own as bugfixes; patch 3 depends on them. For the series, Acked-by: Tejun Heo <[email protected]> Thanks. -- tejun