Re: [PATCH RFC 07/14] fs/erofs: mm/pagemap: add readahead_folio_reverse() to avoid folio->private

"Zi Yan" <[email protected]>
Newsgroups org.kernel.vger.linux-fsdevel,org.kernel.vger.linux-kernel,org.kvack.linux-mm,org.ozlabs.lists.linux-erofs
Message-ID <[email protected]>
On Tue Aug 4, 2026 at 1:09 PM EDT, Zi Yan wrote:
> On Tue Aug 4, 2026 at 1:04 PM EDT, Jan Kara wrote:
>> On Tue 04-08-26 11:54:41, Zi Yan wrote:
>>> On Tue Aug 4, 2026 at 5:32 AM EDT, Jan Kara wrote:
>>> > On Mon 03-08-26 12:56:36, Zi Yan wrote:
>>> >> On Mon Aug 3, 2026 at 5:54 AM EDT, Jan Kara wrote:
>>> >> > On Fri 31-07-26 22:13:30, Zi Yan wrote:
>>> >> >> erofs needs to traverse readahead folios in reverse order to achieve
>>> >> >> maximum performance by
>>> >> >> 1. reading all folios from readahead_folio();
>>> >> >> 2. storing the prior folio pointer in folio->private;
>>> >> >> 3. traverse from the last folio to the first one.
>>> >> >> 
>>> >> >> Add readahead_folio_reverse() to achieve the same function without using
>>> >> >> folio->private.
>>> >> >> 
>>> >> >> It prepares for a future commit that replaces PG_private checks with
>>> >> >> !folio->private checks. After switching the checks, erofs's use of
>>> >> >> folio->private without bumping folio refcount can cause unexpected
>>> >> >> outcomes, e.g., in filemap_release_folio(), try_to_free_buffers() becomes
>>> >> >> reachable.
>>> 
>>> <snip>
>>> 
>>> >> 
>>> >> The below is what I come up with. I did not add a bool to
>>> >> readahead_control, since I think that is the decision of caller of
>>> >> __readahead_advance(). But let me know if you disagree.
>>> >
>>> > The reason why I wanted bool in readahead_control is that if some code
>>> > ends up mixing readahead_folio() with readahead_folio_last() things will
>>> > get confused (because __readahead_advance() really wants to skip the batch
>>> > returned from the *previous* call to readahead_folio[_last]()). With the
>>> > bool in rac, even mixed use will properly advance the state of the
>>> > readahead_control. I don't think mixed use is very realistic (at this
>>> > point at least) so I'm ok with leaving that for later if you don't like it.
>>> 
>>> Got it. I am trying to figure out your mental model of how the mix of
>>> readahead_folio() and readahead_folio_last() works with the bool inside
>>> ractl. By looking at readahead_folio_last() code, it is almost the same
>>> as readahead_folio() with __readahead_folio() inlined
>>> (__readahead_folio() is only used by readahead_folio(), so the inline
>>> can happen without any issue). As a result, we can get rid of
>>> readahead_folio_last(), add set_readahead_direction() to set the
>>> embedded bool read_from_head, and use readahead_folio() only. This
>>> removes redundant code in readahead_folio_last(). One thing I am not
>>> certain is whether we want to
>>> 
>>> 1. use set_readahead_direction() explicit and warn readahead_folio() if
>>> read_from_head is not initialized, or
>>> 
>>> 2. set read_from_head to true by default, so that only erofs needs to
>>> call set_readahead_direction() to change read_from_head.
>>> 
>>> The former is less confusing but changes how readahead_folio() works;
>>> the latter is simpler but implicit read_from_head state might confuse
>>> people at some point.
>>
>> My idea was: readahead_folio() will call __readahead_advance() and then set
>> rac->forward = true. readahead_folio_last() will call __readahead_advance()
>> and set rac->forward = false. __readahead_advance() advances from beginning
>> / end based on rac->_forward value.
>
> Got it. I can do that. Just to be clear, it should be that
> readahead_folio() first sets rac->forward = true, then calls
> __readahead_advance(), since __readahead_advance() advances based on
> rac->forward, right? readahead_folio_last() as well.

This is revised patch:


diff --git a/fs/erofs/zdata.c b/fs/erofs/zdata.c
index 74520e9102596..23f423c22ac8c 100644
--- a/fs/erofs/zdata.c
+++ b/fs/erofs/zdata.c
@@ -1902,21 +1902,14 @@ static void z_erofs_readahead(struct readahead_control *rac)
 	struct inode *realinode = erofs_real_inode(sharedinode, &need_iput);
 	Z_EROFS_DEFINE_FRONTEND(f, realinode, sharedinode, readahead_pos(rac));
 	unsigned int nrpages = readahead_count(rac);
-	struct folio *head = NULL, *folio;
+	struct folio *folio;
 	int err;
 
 	trace_erofs_readahead(realinode, readahead_index(rac), nrpages, false);
 	z_erofs_pcluster_readmore(&f, rac, true);
-	while ((folio = readahead_folio(rac))) {
-		folio->private = head;
-		head = folio;
-	}
-
-	/* traverse in reverse order for best metadata I/O performance */
-	while (head) {
-		folio = head;
-		head = folio_get_private(folio);
 
+	/* traverse from last to first for best metadata I/O performance */
+	while ((folio = readahead_folio_last(rac))) {
 		err = z_erofs_scan_folio(&f, folio, true);
 		if (err && err != -EINTR)
 			erofs_err(realinode->i_sb, "readahead error at folio %lu @ nid %llu",
diff --git a/include/linux/pagemap.h b/include/linux/pagemap.h
index 4e8b2b29f6d3e..cc69d60b2a9d2 100644
--- a/include/linux/pagemap.h
+++ b/include/linux/pagemap.h
@@ -1448,6 +1448,7 @@ struct readahead_control {
 	bool dropbehind;
 	bool _workingset;
 	unsigned long _pflags;
+	bool forward;
 };
 
 #define DEFINE_READAHEAD(ractl, f, r, m, i)				\
@@ -1512,18 +1513,25 @@ void page_cache_async_readahead(struct address_space *mapping,
 	page_cache_async_ra(&ractl, folio, req_count);
 }
 
+static inline void __readahead_advance(struct readahead_control *rac)
+{
+	if (rac->forward)
+		rac->_index += rac->_batch_count;
+
+	rac->_nr_pages -= rac->_batch_count;
+	rac->_batch_count = 0;
+}
+
 static inline struct folio *__readahead_folio(struct readahead_control *ractl)
 {
 	struct folio *folio;
 
 	BUG_ON(ractl->_batch_count > ractl->_nr_pages);
-	ractl->_nr_pages -= ractl->_batch_count;
-	ractl->_index += ractl->_batch_count;
+	ractl->forward = true;
+	__readahead_advance(ractl);
 
-	if (!ractl->_nr_pages) {
-		ractl->_batch_count = 0;
+	if (!ractl->_nr_pages)
 		return NULL;
-	}
 
 	folio = xa_load(&ractl->mapping->i_pages, ractl->_index);
 	VM_BUG_ON_FOLIO(!folio_test_locked(folio), folio);
@@ -1549,6 +1557,39 @@ static inline struct folio *readahead_folio(struct readahead_control *ractl)
 	return folio;
 }
 
+/**
+ * readahead_folio_last - Get the next folio to read, from the tail.
+ * @ractl: The current readahead request.
+ *
+ * Like readahead_folio(), but walks the range back-to-front. The folio is
+ * returned locked with its refcount dropped; the caller unlocks it once I/O
+ * completes. Compound folios are returned once, at their head index.
+ *
+ * Context: The folio is locked.
+ * Return: A pointer to the next folio, or %NULL when done.
+ */
+static inline struct folio *readahead_folio_last(struct readahead_control *ractl)
+{
+	struct folio *folio;
+
+	/* Shrink the window from the tail down to this folio's head index */
+	ractl->forward = false;
+	__readahead_advance(ractl);
+
+	if (!ractl->_nr_pages)
+		return NULL;
+
+	/* xa_load() follows sibling entries, so a tail index returns the head */
+	folio = xa_load(&ractl->mapping->i_pages,
+			ractl->_index + ractl->_nr_pages - 1);
+	VM_WARN_ON_ONCE_FOLIO(!folio_test_locked(folio), folio);
+
+	ractl->_batch_count = folio_nr_pages(folio);
+
+	folio_put(folio);
+	return folio;
+}
+
 static inline unsigned int __readahead_batch(struct readahead_control *rac,
 		struct page **array, unsigned int array_sz)
 {
@@ -1557,9 +1598,8 @@ static inline unsigned int __readahead_batch(struct readahead_control *rac,
 	struct folio *folio;
 
 	BUG_ON(rac->_batch_count > rac->_nr_pages);
-	rac->_nr_pages -= rac->_batch_count;
-	rac->_index += rac->_batch_count;
-	rac->_batch_count = 0;
+	rac->forward = true;
+	__readahead_advance(rac);
 
 	xas_set(&xas, rac->_index);
 	rcu_read_lock();


-- 
Best Regards,
Yan, Zi
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.