Re: [PATCH v7 4/4] perf sched latency: Add histogram and time interval options

Namhyung Kim <[email protected]>
Newsgroups gmane.linux.kernel,gmane.linux.kernel.perf.user
Message-ID <[email protected]>
On Sun, Aug 02, 2026 at 05:09:14PM -0400, Aaron Tomlin wrote:
> While 'perf sched latency' reports task runtime and delay statistics
> (average and maximum delay), it does not provide a visual representation
> of how task wait times are distributed across latency ranges between
> snapshots (start and finish of the analysis window).
> 
> The --histogram option collects CPU wait latencies (time between when
> a task becomes runnable and when it gets scheduled onto a CPU) into 22
> latency buckets, displaying an ASCII bar chart distribution.
> 
> The --hist-mode option configures the bucketing scheme:
>   - log (default). Logarithmic latency buckets ranging from
>     sub-microsecond (< 1 us) up to >= 1.05 seconds
> 
>   - linear. Equal-width linear latency buckets
>     (i.e., 100 us steps up to >= 2.1 ms)
> 
> The --time option allows filtering trace event processing to a
> specific time interval [start,stop].
> 
> Example histogram output excerpt:
> 
>     ❯ sudo perf sched latency --histogram --CPU 0
> 
>      CPU Wait Latency Distribution Histogram (between snapshots) (total samples: 36114)
>      -------------------------------------------------------------------
>       Latency Range    |      Count |    Pct | Histogram Graph
>      -------------------------------------------------------------------
>       < 1 us           |         17 |   0.0% | #
>       2 - 4 us         |        673 |   1.9% | #
>       4 - 8 us         |       6237 |  17.3% | ######
>       8 - 16 us        |       3224 |   8.9% | ###
>       16 - 32 us       |       1388 |   3.8% | #
>       32 - 64 us       |        709 |   2.0% | #
>       64 - 128 us      |        690 |   1.9% | #
>       128 - 256 us     |        789 |   2.2% | #
>       256 - 512 us     |        541 |   1.5% | #
>       512 - 1024 us    |       2256 |   6.2% | ##
>       1 - 2 ms         |       3577 |   9.9% | ###
>       2 - 4 ms         |      13259 |  36.7% | ##############
>       4 - 8 ms         |       2523 |   7.0% | ##
>       8 - 16 ms        |        222 |   0.6% | #
>       16 - 32 ms       |         10 |   0.0% | #
>       >= 1.05 s        |          3 |   0.0% | #
>      -------------------------------------------------------------------
> 
> Signed-off-by: Aaron Tomlin <[email protected]>
> ---
[SNIP] 
> @@ -1168,7 +1306,13 @@ add_sched_in_event(struct work_atoms *atoms, u64 timestamp)
>  		atoms->max_lat_start = atom->wake_up_time;
>  		atoms->max_lat_end = timestamp;
>  	}
> +
>  	atoms->nb_atoms++;
> +
> +	b = latency_bucket(sched, delta);
> +	atoms->hist[b]++;
> +	if (strcmp(thread__comm_str(atoms->thread), "swapper"))
> +		sched->global_hist[b]++;

Why is the swapper thread not included in the global hist?

Also it's probably better to check thread__tid being 0.

>  }
>  
>  static void free_work_atoms(struct work_atoms *atoms)
[SNIP] 
> @@ -3659,6 +3831,21 @@ static int perf_sched__lat(struct perf_sched *sched)
>  	perf_sched__merge_lat(sched);
>  	perf_sched__sort_lat(sched);
>  
> +	next = rb_first_cached(&sched->sorted_atom_root);
> +	while (next) {
> +		struct work_atoms *work_list = rb_entry(next, struct work_atoms, node);
> +
> +		if (work_list->nb_atoms && strcmp(thread__comm_str(work_list->thread), "swapper"))

Ditto.  Comparing TID would be faster.

Thanks,
Namhyung


> +			break;
> +		next = rb_next(next);
> +	}
> +
> +	if (!next) {
> +		pr_info("No matching trace samples found.\n");
> +		rc = 0;
> +		goto out_free_atoms;
> +	}
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.