Re: [PATCH RFC 07/14] fs/erofs: mm/pagemap: add readahead_folio_reverse() to avoid folio->private

Jan Kara <[email protected]>
Newsgroups gmane.linux.kernel,gmane.linux.kernel.mm,gmane.linux.file-systems
Message-ID <2evaxdu6cobpnzzer3y7fsrqvmtmhj7gm3e5buebdaw564igx6@7bs5enipxnpx>
On Tue 04-08-26 11:54:41, Zi Yan wrote:
> On Tue Aug 4, 2026 at 5:32 AM EDT, Jan Kara wrote:
> > On Mon 03-08-26 12:56:36, Zi Yan wrote:
> >> On Mon Aug 3, 2026 at 5:54 AM EDT, Jan Kara wrote:
> >> > On Fri 31-07-26 22:13:30, Zi Yan wrote:
> >> >> erofs needs to traverse readahead folios in reverse order to achieve
> >> >> maximum performance by
> >> >> 1. reading all folios from readahead_folio();
> >> >> 2. storing the prior folio pointer in folio->private;
> >> >> 3. traverse from the last folio to the first one.
> >> >> 
> >> >> Add readahead_folio_reverse() to achieve the same function without using
> >> >> folio->private.
> >> >> 
> >> >> It prepares for a future commit that replaces PG_private checks with
> >> >> !folio->private checks. After switching the checks, erofs's use of
> >> >> folio->private without bumping folio refcount can cause unexpected
> >> >> outcomes, e.g., in filemap_release_folio(), try_to_free_buffers() becomes
> >> >> reachable.
> 
> <snip>
> 
> >> 
> >> The below is what I come up with. I did not add a bool to
> >> readahead_control, since I think that is the decision of caller of
> >> __readahead_advance(). But let me know if you disagree.
> >
> > The reason why I wanted bool in readahead_control is that if some code
> > ends up mixing readahead_folio() with readahead_folio_last() things will
> > get confused (because __readahead_advance() really wants to skip the batch
> > returned from the *previous* call to readahead_folio[_last]()). With the
> > bool in rac, even mixed use will properly advance the state of the
> > readahead_control. I don't think mixed use is very realistic (at this
> > point at least) so I'm ok with leaving that for later if you don't like it.
> 
> Got it. I am trying to figure out your mental model of how the mix of
> readahead_folio() and readahead_folio_last() works with the bool inside
> ractl. By looking at readahead_folio_last() code, it is almost the same
> as readahead_folio() with __readahead_folio() inlined
> (__readahead_folio() is only used by readahead_folio(), so the inline
> can happen without any issue). As a result, we can get rid of
> readahead_folio_last(), add set_readahead_direction() to set the
> embedded bool read_from_head, and use readahead_folio() only. This
> removes redundant code in readahead_folio_last(). One thing I am not
> certain is whether we want to
> 
> 1. use set_readahead_direction() explicit and warn readahead_folio() if
> read_from_head is not initialized, or
> 
> 2. set read_from_head to true by default, so that only erofs needs to
> call set_readahead_direction() to change read_from_head.
> 
> The former is less confusing but changes how readahead_folio() works;
> the latter is simpler but implicit read_from_head state might confuse
> people at some point.

My idea was: readahead_folio() will call __readahead_advance() and then set
rac->forward = true. readahead_folio_last() will call __readahead_advance()
and set rac->forward = false. __readahead_advance() advances from beginning
/ end based on rac->_forward value.

								Honza


> > Also I have some minor comments below.
> >
> >> diff --git a/fs/erofs/zdata.c b/fs/erofs/zdata.c
> >> index b59f2745a8e72..23f423c22ac8c 100644
> >> --- a/fs/erofs/zdata.c
> >> +++ b/fs/erofs/zdata.c
> >> @@ -1908,8 +1908,8 @@ static void z_erofs_readahead(struct readahead_control *rac)
> >>  	trace_erofs_readahead(realinode, readahead_index(rac), nrpages, false);
> >>  	z_erofs_pcluster_readmore(&f, rac, true);
> >>  
> >> -	/* traverse in reverse order for best metadata I/O performance */
> >> -	while ((folio = readahead_folio_reverse(rac))) {
> >> +	/* traverse from last to first for best metadata I/O performance */
> >> +	while ((folio = readahead_folio_last(rac))) {
> >>  		err = z_erofs_scan_folio(&f, folio, true);
> >>  		if (err && err != -EINTR)
> >>  			erofs_err(realinode->i_sb, "readahead error at folio %lu @ nid %llu",
> >> diff --git a/include/linux/pagemap.h b/include/linux/pagemap.h
> >> index 5ca5aa365f319..2cc3de5594518 100644
> >> --- a/include/linux/pagemap.h
> >> +++ b/include/linux/pagemap.h
> >> @@ -1510,13 +1510,21 @@ void page_cache_async_readahead(struct address_space *mapping,
> >>  	page_cache_async_ra(&ractl, folio, req_count);
> >>  }
> >>  
> >> +static inline void __readahead_advance(struct readahead_control *rac,
> >> +		bool read_from_head)
> >> +{
> >> +	if (read_from_head)
> >> +		rac->_index += rac->_batch_count;
> >> +
> >> +	rac->_nr_pages -= rac->_batch_count;
> >> +}
> >
> > Maybe we can add:
> >
> > 	rac->_batch_count = 0;
> >
> > as well since after the advance the _batch_count isn't valid anymore? We
> > can then also remove it from the callers.
> 
> Yes, for __readahead_batch(), _batch_count is zeroed right after. For
> __readahead_folio() and readahead_folio_last(), _batch_count is either
> zeroed or overwritten by folio_nr_pages() before any use.
> 
> >
> >> +
> >>  static inline struct folio *__readahead_folio(struct readahead_control *ractl)
> >>  {
> >> -	struct folio *folio;
> >> +	struct folio *folio = NULL;
> >
> > Not sure why this initialization got here...
> 
> Will remove it. It is some leftover during my development. Thank you for
> pointing it out.
> 
> >
> >>  
> >>  	BUG_ON(ractl->_batch_count > ractl->_nr_pages);
> >> -	ractl->_nr_pages -= ractl->_batch_count;
> >> -	ractl->_index += ractl->_batch_count;
> >> +	__readahead_advance(ractl, /* read_from_head= */ true);
> > 				   ^^^^
> > This is not really a kernel style :), please delete this comment. I'm ok with
> > pure false/true here - it is an internal helper used in few places. If
> > things get wider use, we tend to switch to 'unsigned flags' with explicit
> > flag names to ease code reading. But that's not the case here.
> 
> We did this in some MM code. I will delete the comment like you
> suggested.
> 
> <snip>
> 
> >> @@ -1583,11 +1594,10 @@ static inline unsigned int __readahead_batch(struct readahead_control *rac,
> >>  {
> >>  	unsigned int i = 0;
> >>  	XA_STATE(xas, &rac->mapping->i_pages, 0);
> >> -	struct folio *folio;
> >> +	struct folio *folio = NULL;
> >
> > Again not sure why this initialization got here...
> 
> Will remove.
> 
> -- 
> Best Regards,
> Yan, Zi
> 
-- 
Jan Kara <[email protected]>
SUSE Labs, CR
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.