Re: [PATCH v3 1/1] pipe: only enable the extra wake_up(rd_wait) for edge-triggered consumers

Oleg Nesterov <[email protected]> Fri, 31 Jul 2026 12:45:40 +0200
Newsgroups org.kernel.vger.io-uring,org.kernel.vger.linux-fsdevel,org.kernel.vger.linux-kernel
Message-ID <[email protected]>
On 07/30, Linus Torvalds wrote:
>
> On Thu, 30 Jul 2026 at 13:35, Oleg Nesterov <[email protected]> wrote:
> >
> > Again, contrary to the current comments this is not Epoll-only...
>
> Sure it is. As far as we know, only epoll has ever *cared*.
>
> Yes, you can show the semantics with io_uring, but do you have a user
> application that actually cares?

No. But I know nothing about io_uring, this is the question for Pavel
and Jens.

io_uring claims itself edge-triggered (whatever that means). It has
IORING_POLL_ADD_LEVEL, but this mode was disabled by d59bd748db0a9
("io_uring/poll: disable level triggered poll").

I don't understand io_uring/poll.c even remotely, but it seems that
without IORING_POLL_ADD_LEVEL io_uring expects that the io_poll_wake()
callback should be called on every write.

Same for epoll(EPOLLET)... I mean, I have no idea why anyone would
need this behavior. But from the commit 3a34b13a88caeb28 ("pipe: make
pipe writes always wake up readers") we know that such users exist.

> The only reason that ugly hack exists is because we did have that
> break user space. If we can get rid of th eugly hack for io_uring,
> that would only be a good thing.
>
> So this literally *should* be about only epoll unless you have a
> report that io_uring users are equally broken and use that
> shit-for-brains notion of edges that aren't edges that nobody sane
> should ever use.

In the 1st version I did:

	static void pipe_set_epoll_usage(struct file *filp, struct pipe_inode_info *pipe)
	{
	#ifdef CONFIG_EPOLL
		if ((filp->f_mode & FMODE_READ) && filp->f_ep &&
		   unlikely(!READ_ONCE(pipe->epoll_usage)))
		       WRITE_ONCE(pipe->epoll_usage, true);
	#endif
	}

but sashiko didn't like it, let me quote:

	Does skipping this wakeup for non-epoll consumers break io_uring?

        Applications polling pipes via io_uring do not attach an eventpoll context,
        so pipe->epoll_usage will be false. If a writer writes to an empty pipe,
        io_uring receives the wakeup.

        If the writer then writes a second chunk before the first is drained,
        anon_pipe_write() observes was_empty == false and
        pipe_get_epoll_usage() == false, skipping the waitqueue wakeup.

        Could this cause io_uring to miss events and hang permanently, waiting
        for a CQE that will never be emitted for the new data?

See https://lore.kernel.org/all/[email protected]/ for details.

So I wrote that "Test-case for io_uring" to a) check that sashiko was right,
and b) to ensure that V2 doesn't change the current behaviour of io_uring.

I won't argue with "nobody sane should ever use", but IMO the same is true
for epoll with EPOLLET.

--------------------------------------------------------------------------
So, let me ask. Apart from the comments and naming, do you agree with this
patch?

I like the new version more, even if we forget about io_uring. Note that
this way epoll_ctl() without EPOLLET in .events will not set ->poll_usage,
and hopefully "nobody sane" use this flag...

What do you think?

Oleg.