[PATCH v6] perf record: Fix unhandled POLLHUP on non_perf_event descriptors

Ian Rogers <[email protected]>
Newsgroups org.kernel.vger.linux-perf-users,org.kernel.vger.linux-kernel
Message-ID <[email protected]>
When processing POLLHUP on a non_perf_event control descriptor (like ctl_fd),
evlist__finalize_ctlfd() is invoked to finalize the setup, correctly setting
the core evlist's poll array file descriptor to -1. However, this finalized
teardown state is never propagated back to the individual thread's local
replica of the pollfd array.

Consequently, on the next iteration of the main recording loop,
record__update_evlist_pollfd_from_thread() performs a strict equivalence
check between the core evlist's array and the thread's localized poll array,
detecting that the fd values no longer match (-1 != original_fd). This causes
an immediate -EINVAL abort and a premature, ungraceful teardown.

Fix the underlying logic within record__update_evlist_pollfd_from_thread()
to sustainably propagate the finalized teardown statuses (-1) originating
from the core evlist back to the thread's localized poll structure. This
correctly maintains synchronization and entirely prevents the unhandled
index mismatch crashes.

Additionally, add a unit test that explicitly validates that fdarray__filter()
preserves its invariants regarding fdarray_flag__nonfilterable items to
guard against future regressions.

Fixes: fb4751e79c45 ("perf record: Fix teardown hang on system-wide multi-threaded sessions")
Assisted-by: Gemini:gemini-3.1-pro
Signed-off-by: Ian Rogers <[email protected]>
---
 tools/perf/builtin-record.c | 10 ++++++++++
 tools/perf/tests/fdarray.c  | 24 ++++++++++++++++++++++++
 2 files changed, 34 insertions(+)

diff --git a/tools/perf/builtin-record.c b/tools/perf/builtin-record.c
index a57987851cf0..d7c083803029 100644
--- a/tools/perf/builtin-record.c
+++ b/tools/perf/builtin-record.c
@@ -1169,6 +1169,16 @@ static int record__update_evlist_pollfd_from_thread(struct record *rec,
 		int e_pos = rec->index_map[i].evlist_pollfd_index;
 		int t_pos = rec->index_map[i].thread_pollfd_index;
 
+		if (e_entries[e_pos].fd == -1 || e_entries[e_pos].events == 0) {
+			/*
+			 * e_entries might have been finalized by evlist__finalize_ctlfd().
+			 * We must propagate it to t_entries to avoid index mismatches
+			 * and to prevent a poll storm on the next iteration.
+			 */
+			t_entries[t_pos].fd = -1;
+			t_entries[t_pos].events = 0;
+		}
+
 		if (e_entries[e_pos].fd != t_entries[t_pos].fd ||
 		    e_entries[e_pos].events != t_entries[t_pos].events) {
 			pr_err("Thread and evlist pollfd index mismatch\n");
diff --git a/tools/perf/tests/fdarray.c b/tools/perf/tests/fdarray.c
index 40983c3574b1..2d3db7b754a1 100644
--- a/tools/perf/tests/fdarray.c
+++ b/tools/perf/tests/fdarray.c
@@ -80,6 +80,30 @@ static int test__fdarray__filter(struct test_suite *test __maybe_unused, int sub
 		goto out_delete;
 	}
 
+	fdarray__init_revents(fda, POLLHUP);
+	fda->priv[2].flags = fdarray_flag__nonfilterable;
+
+	pr_debug("\nfiltering all but fda->entries[2] (nonfilterable):");
+	fdarray__fprintf_prefix(fda, "before", stderr);
+	nr_fds = fdarray__filter(fda, POLLHUP, NULL, NULL);
+	fdarray__fprintf_prefix(fda, " after", stderr);
+
+	if (nr_fds != 0) {
+		pr_debug("\nfdarray__filter()=%d != 0, should be 0\n",
+			 nr_fds);
+		goto out_delete;
+	}
+	if (fda->entries[2].fd == -1) {
+		pr_debug("\nfdarray__filter() illegally modified nonfilterable fd!");
+		goto out_delete;
+	}
+	if (fda->entries[2].revents != POLLHUP) {
+		pr_debug("\nfdarray__filter() illegally modified nonfilterable revents!");
+		goto out_delete;
+	}
+
+	fda->priv[2].flags = 0; /* reset flags */
+
 	pr_debug("\n");
 
 	err = 0;
-- 
2.55.0.737.g08866a6d13-goog
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.