Re: [PATCH v7 4/4] perf sched latency: Add histogram and time interval options
Namhyung Kim <[email protected]> Mon, 3 Aug 2026 11:06:16 -0700
| Newsgroups | org.kernel.vger.linux-perf-users,org.kernel.vger.linux-kernel |
|---|---|
| Message-ID | <[email protected]> |
On Sun, Aug 02, 2026 at 05:09:14PM -0400, Aaron Tomlin wrote: > While 'perf sched latency' reports task runtime and delay statistics > (average and maximum delay), it does not provide a visual representation > of how task wait times are distributed across latency ranges between > snapshots (start and finish of the analysis window). > > The --histogram option collects CPU wait latencies (time between when > a task becomes runnable and when it gets scheduled onto a CPU) into 22 > latency buckets, displaying an ASCII bar chart distribution. > > The --hist-mode option configures the bucketing scheme: > - log (default). Logarithmic latency buckets ranging from > sub-microsecond (< 1 us) up to >= 1.05 seconds > > - linear. Equal-width linear latency buckets > (i.e., 100 us steps up to >= 2.1 ms) > > The --time option allows filtering trace event processing to a > specific time interval [start,stop]. > > Example histogram output excerpt: > > ❯ sudo perf sched latency --histogram --CPU 0 > > CPU Wait Latency Distribution Histogram (between snapshots) (total samples: 36114) > ------------------------------------------------------------------- > Latency Range | Count | Pct | Histogram Graph > ------------------------------------------------------------------- > < 1 us | 17 | 0.0% | # > 2 - 4 us | 673 | 1.9% | # > 4 - 8 us | 6237 | 17.3% | ###### > 8 - 16 us | 3224 | 8.9% | ### > 16 - 32 us | 1388 | 3.8% | # > 32 - 64 us | 709 | 2.0% | # > 64 - 128 us | 690 | 1.9% | # > 128 - 256 us | 789 | 2.2% | # > 256 - 512 us | 541 | 1.5% | # > 512 - 1024 us | 2256 | 6.2% | ## > 1 - 2 ms | 3577 | 9.9% | ### > 2 - 4 ms | 13259 | 36.7% | ############## > 4 - 8 ms | 2523 | 7.0% | ## > 8 - 16 ms | 222 | 0.6% | # > 16 - 32 ms | 10 | 0.0% | # > >= 1.05 s | 3 | 0.0% | # > ------------------------------------------------------------------- > > Signed-off-by: Aaron Tomlin <[email protected]> > --- [SNIP] > @@ -1168,7 +1306,13 @@ add_sched_in_event(struct work_atoms *atoms, u64 timestamp) > atoms->max_lat_start = atom->wake_up_time; > atoms->max_lat_end = timestamp; > } > + > atoms->nb_atoms++; > + > + b = latency_bucket(sched, delta); > + atoms->hist[b]++; > + if (strcmp(thread__comm_str(atoms->thread), "swapper")) > + sched->global_hist[b]++; Why is the swapper thread not included in the global hist? Also it's probably better to check thread__tid being 0. > } > > static void free_work_atoms(struct work_atoms *atoms) [SNIP] > @@ -3659,6 +3831,21 @@ static int perf_sched__lat(struct perf_sched *sched) > perf_sched__merge_lat(sched); > perf_sched__sort_lat(sched); > > + next = rb_first_cached(&sched->sorted_atom_root); > + while (next) { > + struct work_atoms *work_list = rb_entry(next, struct work_atoms, node); > + > + if (work_list->nb_atoms && strcmp(thread__comm_str(work_list->thread), "swapper")) Ditto. Comparing TID would be faster. Thanks, Namhyung > + break; > + next = rb_next(next); > + } > + > + if (!next) { > + pr_info("No matching trace samples found.\n"); > + rc = 0; > + goto out_free_atoms; > + }