Re: [PATCH] amdkfd: allow profiler process reenter

"Kuehling, Felix" <[email protected]>
Newsgroups org.freedesktop.lists.amd-gfx
Message-ID <[email protected]>
On 2026-08-13 02:21, Zhu, Lingshan wrote:
> Hello Felix,
>
> Yes I agree the user space may be buggy, and the kernel should be robust.
>
> Even though it is named locking, it is actually a reservation, not recursive locking.
> The kfd_process that successfully invokes the ioctl reserves the device as a profiler.
> So when a user space locks the device multiple times, it does not increase any reference counters, and doesn't need to unlock multiple times.
>
> Multiple times lockings from the same kfd_process just keeps the reservation, it still reserves the device successfully,
> so IMHO kfd should return 0 other than -EALREADY, which is a negative error code.

That's exactly the problem. The API is apparently designed for user mode 
to "reserve" profiler access once, presumably to get exclusive access to 
some profiler resources. If user space reserves multiple times, you 
probably have multiple uncoordinated users of the profiler features in 
the same process. If they are not aware of each other, they will run 
into problems because they're expecting exclusive access to resources 
that are actually shared with someone else.

For example they will both independently call "unreserve" as well. When 
the first one calls "unreserve", the other one breaks. This may happen 
intermittently, depending on the order of operations. Or it may be 
noticed as a regression after seemingly harmless changes that affect the 
order of operations. This is much harder to debug than a clear error 
code returned when a double reservation first happens.

By returning 0 here, you paper over the user mode bug. But chances are, 
you're going to run into other bugs later that have the same cause and 
are much more difficult to debug. Unless someone explains to me that I'm 
misunderstanding the intent of the reservation API, I will not approve 
this change.


>
> By the way, the function is declared to return uint32_t, which is definitely wrong, because the function may return errors, for example -EBUSY.

I agree with this part.

Regards,
   Felix


>
> Thanks
> Lingshan
> On 8/13/2026 12:27 AM, Kuehling, Felix wrote:
>> I suspect this won't work correctly. If the profiler locks the device 
>> multiple times, it probably also unlocks it multiple times. But I see 
>> no reference counting being done, so this may result in a situation 
>> where user mode locks twice, and unlocks once. Now it's lost the lock 
>> but it was still expecting to have it from the first lock operation.
>>
>> I suspect that this is a user mode bug, and KFD is working as expected.
>>
>> Regards,
>>   Felix
>>
>> On 2026-08-12 03:14, Zhu, Lingshan wrote:
>>> AMD General
>>>
>>> Gentle Ping.
>>>
>>> Thanks
>>> Lingshan
>>>
>>> -----Original Message-----
>>> From: Zhu, Lingshan <[email protected]>
>>> Sent: Thursday, August 6, 2026 4:43 PM
>>> To: Deucher, Alexander <[email protected]>; Koenig, 
>>> Christian <[email protected]>; Kuehling, Felix 
>>> <[email protected]>
>>> Cc: Huang, Ray <[email protected]>; [email protected]; 
>>> Zhu, Lingshan <[email protected]>
>>> Subject: [PATCH] amdkfd: allow profiler process reenter
>>>
>>> profile_lock_device() currently returns -EALREADY when the current 
>>> profiler process requests to lock the device again. However any 
>>> negtive values here are treated as failure, so this causes rejecting 
>>> a profiler operation even though the caller still owns the device 
>>> locking.
>>>
>>> This commit fixes this problem by returning 0 when the current 
>>> profiler process locking the device again.
>>>
>>> This commit also changes profile_lock_device() to return an int 
>>> value other than uint32_t, because it returns negative number when 
>>> fail.
>>>
>>> Signed-off-by: Zhu Lingshan <[email protected]>
>>> ---
>>>   drivers/gpu/drm/amd/amdkfd/kfd_chardev.c | 6 +++---
>>>   1 file changed, 3 insertions(+), 3 deletions(-)
>>>
>>> diff --git a/drivers/gpu/drm/amd/amdkfd/kfd_chardev.c 
>>> b/drivers/gpu/drm/amd/amdkfd/kfd_chardev.c
>>> index 309510e23315..9acd5c91e51c 100644
>>> --- a/drivers/gpu/drm/amd/amdkfd/kfd_chardev.c
>>> +++ b/drivers/gpu/drm/amd/amdkfd/kfd_chardev.c
>>> @@ -3329,8 +3329,8 @@ static int kfd_ioctl_create_process(struct 
>>> file *filep, struct kfd_process *p, v
>>>          return 0;
>>>   }
>>>
>>> -static inline uint32_t profile_lock_device(struct kfd_process *p,
>>> -                                          uint32_t gpu_id, uint32_t 
>>> op)
>>> +static inline int profile_lock_device(struct kfd_process *p,
>>> +                                     uint32_t gpu_id, uint32_t op)
>>>   {
>>>          struct kfd_process_device *pdd;
>>>          struct kfd_dev *kfd;
>>> @@ -3365,7 +3365,7 @@ static inline uint32_t 
>>> profile_lock_device(struct kfd_process *p,
>>>                          }
>>>                          return status;
>>>                  } else if (kfd->profiler_process == p) {
>>> -                       status = -EALREADY;
>>> +                       status = 0;
>>>                  } else {
>>>                          status = -EBUSY;
>>>                  }
>>> -- 
>>> 2.53.0
>>>
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.