Re: [PATCH v4 05/17] irq & spin_lock: Add counted interrupt disabling/enabling

Shrikanth Hegde <[email protected]> Wed, 5 Aug 2026 22:23:20 +0530
Newsgroups org.kernel.vger.rust-for-linux,org.kernel.vger.linux-kernel
Message-ID <[email protected]>
Hi Boqun.

On 8/5/26 8:41 PM, Boqun Feng wrote:
> On Wed, Aug 05, 2026 at 08:26:49PM +0530, Shrikanth Hegde wrote:
>>
>>
>> On 8/5/26 7:50 PM, Boqun Feng wrote:
>>> On Wed, Aug 05, 2026 at 07:40:18PM +0530, Shrikanth Hegde wrote:
>>>>
>>>> Hi Boqun,
>>>>
>>>>>
>>>>> Something as below? Going to send it to kernel build bot and see if it
>>>>> works for all configs.
>>>>>
>>>>> ----------------->8
>>>>> diff --git a/include/linux/interrupt_rc.h b/include/linux/interrupt_rc.h
>>>>> index b9a7f05ecf42..39f30bc65548 100644
>>>>> --- a/include/linux/interrupt_rc.h
>>>>> +++ b/include/linux/interrupt_rc.h
>>>>> @@ -12,6 +12,7 @@
>>>>>      */
>>>>>
>>>>>     #include <linux/irqflags.h>
>>>>> +#include <linux/debug_locks.h>
>>>>>     #include <linux/preempt.h>
>>>>>     #include <linux/processor.h>
>>>>>     #include <linux/smp.h>
>>>>> @@ -63,6 +64,12 @@ static inline void local_interrupt_disable(void)
>>>>>
>>>>>            new_count = hardirq_disable_enter();
>>>>>
>>>>> +       /* Is hardirq disable count overflow soon? */
>>>>> +       if (IS_ENABLED(CONFIG_DEBUG_PREEMPT))
>>>>> +               DEBUG_LOCKS_WARN_ON((new_count & HARDIRQ_DISABLE_MASK) +
>>>>> +                                   (10 << HARDIRQ_DISABLE_SHIFT) >
>>>>> +                                   HARDIRQ_DISABLE_MASK);
>>>>> +
>>>>
>>>> This needs a return here right? Else we will see warning for 10 times
>>>> and then overflow happens and we will call _local_interrupt_disable. No?
> 
> Oh, seems I overlooked something... could you elaborate on this? What's
> the scenario in your mind? You said we hit 10 times warning and *then*
> overflow?
> 


(the “10 times” wording was for the possible wraparound.
I didn't know DEBUG_LOCKS_WARN_ON() will turn debug_locks
off after the first warning, I thought it will print 10 times.)

Main concern is the wraparound itself. If the HARDIRQ_DISABLE field reaches
0xff and we increment it once more, the field becomes zero after masking
(ff + 1 -> 100 in the shifted field). Then local_interrupt_enable() can
observe:

         (new_count & HARDIRQ_DISABLE_MASK) == 0

and treat it as the outermost enable, potentially calling
_local_interrupt_enable() while there are still outstanding logical
local_interrupt_disable() users.

So I think the useful debug check is one that catches the count before it
gets close enough to wrap. Thing I wanted to avoid is silently making the
HARDIRQ_DISABLE field look like zero after overflow.

Whether it returns or just warns, I am not sure. As recovery may
not be easy. It is meant more to be a damage contol than recovery.


My line of thought is,

If we return after debug checks in local_interrupt_disable(), we wont advance the preempt count
further, so it wont wrap around. Now, there will be corresponding local_interrupt_enable(),
at some point it will reach 0, we enable the interrupts. There will be some more
local_interrupt_enable() still, but they will be caught with your debug check in
local_interrupt_enable() and preempt count won't be decremented further.
So if we put return we may have a damage control.

>>>>
>>>
>>> DEBUG_LOCKS_WARN_ON() uses debug_locks_off() to avoid this, so we won't
>>> see it 10 times. The reason not using return here, because we would
>>> introduce unpaired local_interrupt_disable() if we returned:
>>
>> Yes, it could be a weird case if the overflow actually happens.
>> So just warning maybe enough to catch such callers.
>>
>> Maybe your kunit test can actually help test the behavior with the loop
>> count.
>>
>>>
>>> 	// hardirq disable count is n
>>> 	local_interrupt_disable(); // hardirq disable count is n + 1
>>>
>>> 	local_interrupt_disable(); <- trigger the warning, if we return
>>> 				   // hardirq disable count is n + 1
>>
>> Likely i am missing to understand.
>>
> 
> No, it was me who misunderstand ;-)
> 
>> Isn't the count incremented earlier than return?
>> I.e even if return happens it should be n + 2 right?
>>
> 
> Yeah, you're right, but then why do we want to return earlier? Since the
> following if:
> 
> 	if ((new_count & HARDIRQ_DISABLE_MASK) == HARDIRQ_DISABLE_OFFSET)
> 
> will be false, and we will just return from the function, no?
> 
> Regards,
> Boqun
> 
>>>
>>> 	local_interrupt_enable(); // hardirq disable count is n
>>>
>>> 	local_interrupt_enable(); // hardirq disable count is n - 1
>>>
>>> Regards,
>>> Boqun
>>>
>>>> Not sure, if below is any better? (Igore whitespace mangling)
>>>>
>>>> if (IS_ENABLED(CONFIG_DEBUG_PREEMPT) &&
>>>>       DEBUG_LOCKS_WARN_ON((preempt_count() & HARDIRQ_DISABLE_MASK) >=
>>>> 			 HARDIRQ_DISABLE_MASK - (10 << HARDIRQ_DISABLE_SHIFT)))
>>>> 	return;
>>>>
>>>>>            /* Interrupts can happen here, but it's OK, see __irq_exit_rcu(). */
>>>>>
>>>>>            if ((new_count & HARDIRQ_DISABLE_MASK) == HARDIRQ_DISABLE_OFFSET)
>>>>> @@ -73,6 +80,11 @@ static inline void local_interrupt_enable(void)
>>>>>     {
>>>>>            int new_count;
>>>>>
>>>>> +       /* Unpaired local_interrupt_enable()? Warn and abort. */
>>>>> +       if (IS_ENABLED(CONFIG_DEBUG_PREEMPT) &&
>>>>> +           DEBUG_LOCKS_WARN_ON((preempt_count() & HARDIRQ_DISABLE_MASK) == 0))
>>>>> +               return;
>>>>> +
>>>>>            new_count = hardirq_disable_exit();
>>>>>
>>>>>            if ((new_count & HARDIRQ_DISABLE_MASK) == 0) >
>>>>
>>