Re: Removing ->dirty_folio

Matthew Wilcox <[email protected]>
Newsgroups org.kvack.linux-mm,org.kernel.vger.linux-btrfs,org.kernel.vger.linux-fsdevel,org.kernel.vger.linux-xfs
Message-ID <[email protected]>
On Mon, Aug 24, 2026 at 12:51:43PM -0700, John Hubbard wrote:
> On 8/24/26 12:43 PM, Matthew Wilcox wrote:
> > On Mon, Aug 24, 2026 at 12:25:42PM -0700, John Hubbard wrote:
> >>> My proposal is this:
> >>>
> >>>  - Fileystems take note of folio_maybe_dma_pinned() during writeback.
> >>>    If it's true, do the writeback, but retain/recreate whatever data
> >>>    structures you need in order to write the folio again; behave as if
> >>>    ->page_mkdirty() had been called again for each page in the folio is
> >>>    marked as dirty.
> >>
> >> Yes, that would work nicely.
> >>
> >>>  - The MM behaves similarly; we do not clear the writeback flag for
> >>>    folio_maybe_dma_pinned().
> >>>
> >>> This will have the effect of writing pinned folios back every time the
> >>> inode is scheduled for writeback.  But since we have no idea whether
> >>> the folio is actually dirty (because the GUP user won't tell us),
> >>> this is the correct behaviour.
> >>>
> >>> I'm probably missing some stuff here.  Let me know.
> >>
> >> OK, so working through the end of the pinning, I think it still is
> >> correct: device finishes writing to pinned memory, device driver
> >> unpins the memory but the page has been left marked dirty the whole
> 
> oh, I just thought of a minor hole that we need to fill: how to mark the
> page dirty in the first place, in the absence of mark_[page|folio]_dirty()?
> 
> Under this new scheme, we will need to pin first, then mark dirty, to
> set up. The filesystem can't do everything, because even if it were to
> call page_mkdirty(), a writeback could clear that before the page gets
> pinned.

The page fault path:

handle_mm_fault()
  __handle_mm_fault()
    handle_pte_fault()
      do_wp_page() [just assuming the pte is present, but !writable]
        wp_page_shared()
	  do_page_mkwrite()
	    vmf->vma->vm_ops->page_mkwrite(vmf)

and that's where the filesystem gets notified that this page is about
to become writable.

So your concern is obviously "how do we prevent the writeout from
happening before we set the pincount", and I think it's that the
writeout path will make the PTE read-only before it does writeback,
and we hold the mmap_lock which prevents the page table entry from being 
made read-only.

But I'm only about 80% sure that's what happens.

> > @@ -2717,7 +2719,8 @@ static inline bool folio_maybe_dma_pinned(struct folio *folio)
> >          * Here, for that overflow case, use the sign bit to count a little
> >          * bit higher via unsigned math, and thus still get an accurate result.
> >          */
> > -       return ((unsigned int)folio_ref_count(folio)) >=
> > +       mapcount = folio_mapcount(folio);
> > +       return (folio_ref_count(folio) - mapcount) >=
> 
> As long as the math works: need to not underflow. I guess mapcount is
> always less than refcount, so OK.
> 
> So it *seems* correct to me, fwiw. :)

Yeah, and if we hold the folio locked, it's true.  But we could sample
mapcount, then have a few unmaps come in before we read refcount, and
we've got an underflow.  I mean, it's only "maybe" mapped ... ;-)

Perhaps we could have two functions, one for if you have the folio
locked (like in the writeback path) where you can rely on mapcount
not changing, and thus refcount always being > mapcount.
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.