Re: [PATCH v2 3/3] drm/panthor: Add tracepoints for cache flushing

Nicolas Frattaroli <[email protected]> Tue, 04 Aug 2026 16:44:40 +0200
Newsgroups gmane.linux.kernel,gmane.comp.video.dri.devel
Message-ID <[email protected]>
On Monday, 3 August 2026 11:03:58 Central European Summer Time Boris Brezillon wrote:
> On Thu, 30 Jul 2026 13:45:16 +0200
> Nicolas Frattaroli <[email protected]> wrote:
> 
> > Add two new event tracepoints: gpu_cache_flush_start to be emitted after
> > acquiring the flush mutex and reqs spinlock, and gpu_cache_flush_end to
> > be emitted when leaving the function.
> > 
> > This allows debugging the duration a flush takes irrespective of initial
> > function entry lock contention by subtracting the start tracepoint's
> > timestamp from the end tracepoint timestamp, and additionally contains
> > information such as which caches were flushed.
> > 
> > Signed-off-by: Nicolas Frattaroli <[email protected]>
> > ---
> >  drivers/gpu/drm/panthor/panthor_gpu.c   |  3 ++
> >  drivers/gpu/drm/panthor/panthor_trace.h | 49 +++++++++++++++++++++++++++++++++
> >  2 files changed, 52 insertions(+)
> > 
> > diff --git a/drivers/gpu/drm/panthor/panthor_gpu.c b/drivers/gpu/drm/panthor/panthor_gpu.c
> > index f015bde80abf..a25955ad668b 100644
> > --- a/drivers/gpu/drm/panthor/panthor_gpu.c
> > +++ b/drivers/gpu/drm/panthor/panthor_gpu.c
> > @@ -336,6 +336,7 @@ int panthor_gpu_flush_caches(struct panthor_device *ptdev,
> >  	guard(mutex)(&ptdev->gpu->cache_flush_lock);
> >  
> >  	spin_lock(&ptdev->gpu->reqs_lock);
> > +	trace_gpu_cache_flush_start(ptdev->base.dev, l2, lsc, other);
> >  	if (!(ptdev->gpu->pending_reqs & GPU_IRQ_CLEAN_CACHES_COMPLETED)) {
> >  		ptdev->gpu->pending_reqs |= GPU_IRQ_CLEAN_CACHES_COMPLETED;
> >  		gpu_write(gpu->iomem, GPU_CMD, GPU_FLUSH_CACHES(l2, lsc, other));
> > @@ -345,6 +346,7 @@ int panthor_gpu_flush_caches(struct panthor_device *ptdev,
> >  
> >  	if (ret) {
> >  		spin_unlock(&ptdev->gpu->reqs_lock);
> > +		trace_gpu_cache_flush_end(ptdev->base.dev, l2, lsc, other);
> 
> Should we add a status to the end event, so that faulty flushes can
> be filtered out (those can be immediate, or timeout=100ms depending on
> where the failure happens, but they are not reflecting anything useful,
> and would pollute the stats)? Actually, if what we care about is the
> time it takes to do a flush, do we even need those start/end events,
> can't we pass the time as an argument and do the math in the function
> like we do in the job IRQ handler?

The start/end pattern seems more common across the kernel. Since there
is no doubt as to which start correlates with which end due to the
mutex, I chose to go with that.

> 
> >  		return ret;
> >  	}
> >  
> > @@ -358,6 +360,7 @@ int panthor_gpu_flush_caches(struct panthor_device *ptdev,
> >  			ptdev->gpu->pending_reqs &= ~GPU_IRQ_CLEAN_CACHES_COMPLETED;
> >  	}
> >  	spin_unlock(&ptdev->gpu->reqs_lock);
> > +	trace_gpu_cache_flush_end(ptdev->base.dev, l2, lsc, other);
> >  
> >  	if (ret) {
> >  		panthor_device_schedule_reset(ptdev);
> > diff --git a/drivers/gpu/drm/panthor/panthor_trace.h b/drivers/gpu/drm/panthor/panthor_trace.h
> > index 6ffeb4fe6599..6951b95b1de7 100644
> > --- a/drivers/gpu/drm/panthor/panthor_trace.h
> > +++ b/drivers/gpu/drm/panthor/panthor_trace.h
> > @@ -76,6 +76,55 @@ TRACE_EVENT(gpu_job_irq,
> >  		  __entry->events, __entry->duration_ns)
> >  );
> >  
> > +DECLARE_EVENT_CLASS(gpu_cache_flush_template,
> > +	TP_PROTO(const struct device *dev, u32 l2, u32 lsc, u32 other),
> > +	TP_ARGS(dev, l2, lsc, other),
> > +	TP_STRUCT__entry(
> > +		__string(dev_name, dev_name(dev))
> > +		__field(u32, l2)
> > +		__field(u32, lsc)
> > +		__field(u32, other)
> > +	),
> > +	TP_fast_assign(
> > +		__assign_str(dev_name);
> > +		__entry->l2     = l2;
> > +		__entry->lsc    = lsc;
> > +		__entry->other  = other;
> > +	),
> > +	TP_printk("%s: l2=0x%x lsc=0x%x other=0x%x", __get_str(dev_name),
> > +		  __entry->l2, __entry->lsc, __entry->other)
> > +);
> > +
> > +/**
> > + * gpu_cache_flush_start - called after cache flush locks taken, before flush
> > + * @dev: pointer to the &struct device, for printing the device name
> > + * @l2: "l2" flush flags
> > + * @lsc: "lsc" flush flags
> > + * @other: "other" flush flags
> > + *
> > + * Fires after any initial lock contention around the locks needed for flushing
> > + * caches, but before the actual cache flush is requested.
> > + */
> > +DEFINE_EVENT(gpu_cache_flush_template, gpu_cache_flush_start,
> > +	     TP_PROTO(const struct device *dev, u32 l2, u32 lsc, u32 other),
> > +	     TP_ARGS(dev, l2, lsc, other)
> > +);
> > +
> > +/**
> > + * gpu_cache_flush_end - called after cache flush
> > + * @dev: pointer to the &struct device, for printing the device name
> > + * @l2: "l2" flush flags
> > + * @lsc: "lsc" flush flags
> > + * @other: "other" flush flags
> > + *
> > + * Fires after either the cache flush is complete, or has failed. Can be used
> > + * together with gpu_cache_flush_start to get how long the flush has taken.
> > + */
> > +DEFINE_EVENT(gpu_cache_flush_template, gpu_cache_flush_end,
> > +	     TP_PROTO(const struct device *dev, u32 l2, u32 lsc, u32 other),
> > +	     TP_ARGS(dev, l2, lsc, other)
> > +);
> > +
> >  #endif /* __PANTHOR_TRACE_H__ */
> >  
> >  #undef TRACE_INCLUDE_PATH
> > 
> 
>