Re: [PATCH RFC 07/14] fs/erofs: mm/pagemap: add readahead_folio_reverse() to avoid folio->private

Jan Kara <[email protected]> Tue, 4 Aug 2026 19:04:08 +0200
Newsgroups gmane.linux.kernel,gmane.linux.kernel.mm,gmane.linux.file-systems
Message-ID <2evaxdu6cobpnzzer3y7fsrqvmtmhj7gm3e5buebdaw564igx6@7bs5enipxnpx>
On Tue 04-08-26 11:54:41, Zi Yan wrote:
> On Tue Aug 4, 2026 at 5:32 AM EDT, Jan Kara wrote:
> > On Mon 03-08-26 12:56:36, Zi Yan wrote:
> >> On Mon Aug 3, 2026 at 5:54 AM EDT, Jan Kara wrote:
> >> > On Fri 31-07-26 22:13:30, Zi Yan wrote:
> >> >> erofs needs to traverse readahead folios in reverse order to achieve
> >> >> maximum performance by
> >> >> 1. reading all folios from readahead_folio();
> >> >> 2. storing the prior folio pointer in folio->private;
> >> >> 3. traverse from the last folio to the first one.
> >> >> 
> >> >> Add readahead_folio_reverse() to achieve the same function without using
> >> >> folio->private.
> >> >> 
> >> >> It prepares for a future commit that replaces PG_private checks with
> >> >> !folio->private checks. After switching the checks, erofs's use of
> >> >> folio->private without bumping folio refcount can cause unexpected
> >> >> outcomes, e.g., in filemap_release_folio(), try_to_free_buffers() becomes
> >> >> reachable.
> 
> <snip>
> 
> >> 
> >> The below is what I come up with. I did not add a bool to
> >> readahead_control, since I think that is the decision of caller of
> >> __readahead_advance(). But let me know if you disagree.
> >
> > The reason why I wanted bool in readahead_control is that if some code
> > ends up mixing readahead_folio() with readahead_folio_last() things will
> > get confused (because __readahead_advance() really wants to skip the batch
> > returned from the *previous* call to readahead_folio[_last]()). With the
> > bool in rac, even mixed use will properly advance the state of the
> > readahead_control. I don't think mixed use is very realistic (at this
> > point at least) so I'm ok with leaving that for later if you don't like it.
> 
> Got it. I am trying to figure out your mental model of how the mix of
> readahead_folio() and readahead_folio_last() works with the bool inside
> ractl. By looking at readahead_folio_last() code, it is almost the same
> as readahead_folio() with __readahead_folio() inlined
> (__readahead_folio() is only used by readahead_folio(), so the inline
> can happen without any issue). As a result, we can get rid of
> readahead_folio_last(), add set_readahead_direction() to set the
> embedded bool read_from_head, and use readahead_folio() only. This
> removes redundant code in readahead_folio_last(). One thing I am not
> certain is whether we want to
> 
> 1. use set_readahead_direction() explicit and warn readahead_folio() if
> read_from_head is not initialized, or
> 
> 2. set read_from_head to true by default, so that only erofs needs to
> call set_readahead_direction() to change read_from_head.
> 
> The former is less confusing but changes how readahead_folio() works;
> the latter is simpler but implicit read_from_head state might confuse
> people at some point.

My idea was: readahead_folio() will call __readahead_advance() and then set
rac->forward = true. readahead_folio_last() will call __readahead_advance()
and set rac->forward = false. __readahead_advance() advances from beginning
/ end based on rac->_forward value.

								Honza


> > Also I have some minor comments below.
> >
> >> diff --git a/fs/erofs/zdata.c b/fs/erofs/zdata.c
> >> index b59f2745a8e72..23f423c22ac8c 100644
> >> --- a/fs/erofs/zdata.c
> >> +++ b/fs/erofs/zdata.c
> >> @@ -1908,8 +1908,8 @@ static void z_erofs_readahead(struct readahead_control *rac)
> >>  	trace_erofs_readahead(realinode, readahead_index(rac), nrpages, false);
> >>  	z_erofs_pcluster_readmore(&f, rac, true);
> >>  
> >> -	/* traverse in reverse order for best metadata I/O performance */
> >> -	while ((folio = readahead_folio_reverse(rac))) {
> >> +	/* traverse from last to first for best metadata I/O performance */
> >> +	while ((folio = readahead_folio_last(rac))) {
> >>  		err = z_erofs_scan_folio(&f, folio, true);
> >>  		if (err && err != -EINTR)
> >>  			erofs_err(realinode->i_sb, "readahead error at folio %lu @ nid %llu",
> >> diff --git a/include/linux/pagemap.h b/include/linux/pagemap.h
> >> index 5ca5aa365f319..2cc3de5594518 100644
> >> --- a/include/linux/pagemap.h
> >> +++ b/include/linux/pagemap.h
> >> @@ -1510,13 +1510,21 @@ void page_cache_async_readahead(struct address_space *mapping,
> >>  	page_cache_async_ra(&ractl, folio, req_count);
> >>  }
> >>  
> >> +static inline void __readahead_advance(struct readahead_control *rac,
> >> +		bool read_from_head)
> >> +{
> >> +	if (read_from_head)
> >> +		rac->_index += rac->_batch_count;
> >> +
> >> +	rac->_nr_pages -= rac->_batch_count;
> >> +}
> >
> > Maybe we can add:
> >
> > 	rac->_batch_count = 0;
> >
> > as well since after the advance the _batch_count isn't valid anymore? We
> > can then also remove it from the callers.
> 
> Yes, for __readahead_batch(), _batch_count is zeroed right after. For
> __readahead_folio() and readahead_folio_last(), _batch_count is either
> zeroed or overwritten by folio_nr_pages() before any use.
> 
> >
> >> +
> >>  static inline struct folio *__readahead_folio(struct readahead_control *ractl)
> >>  {
> >> -	struct folio *folio;
> >> +	struct folio *folio = NULL;
> >
> > Not sure why this initialization got here...
> 
> Will remove it. It is some leftover during my development. Thank you for
> pointing it out.
> 
> >
> >>  
> >>  	BUG_ON(ractl->_batch_count > ractl->_nr_pages);
> >> -	ractl->_nr_pages -= ractl->_batch_count;
> >> -	ractl->_index += ractl->_batch_count;
> >> +	__readahead_advance(ractl, /* read_from_head= */ true);
> > 				   ^^^^
> > This is not really a kernel style :), please delete this comment. I'm ok with
> > pure false/true here - it is an internal helper used in few places. If
> > things get wider use, we tend to switch to 'unsigned flags' with explicit
> > flag names to ease code reading. But that's not the case here.
> 
> We did this in some MM code. I will delete the comment like you
> suggested.
> 
> <snip>
> 
> >> @@ -1583,11 +1594,10 @@ static inline unsigned int __readahead_batch(struct readahead_control *rac,
> >>  {
> >>  	unsigned int i = 0;
> >>  	XA_STATE(xas, &rac->mapping->i_pages, 0);
> >> -	struct folio *folio;
> >> +	struct folio *folio = NULL;
> >
> > Again not sure why this initialization got here...
> 
> Will remove.
> 
> -- 
> Best Regards,
> Yan, Zi
> 
-- 
Jan Kara <[email protected]>
SUSE Labs, CR