Re: [PATCH tty v11 1/2] serial: 8250: Switch to nbcon console, take 2

Jon Hunter <[email protected]>
Newsgroups gmane.linux.kernel,gmane.linux.serial,gmane.linux.ports.tegra
Message-ID <[email protected]>
Hi John,

On 22/08/2026 01:32, John Ogness wrote:
> On 2026-08-21, Jon Hunter <[email protected]> wrote:
>> I replaced the WARN with a pr_err() and I see the same issue.
> 
> Thanks for testing. But just to be certain, can you restore the original
> code (with the warning) and apply this horrible hack to force atomic
> printing? I want to be certain that the problem goes away when nbcon
> deferred printing is not active. For thet test you should boot normally
> (without keep_bootcon).
> 
> ===== BEGIN HACK =====
> diff --git a/kernel/printk/internal.h b/kernel/printk/internal.h
> index 85fbf1801cbe0..5806815fa11b2 100644
> --- a/kernel/printk/internal.h
> +++ b/kernel/printk/internal.h
> @@ -202,6 +202,7 @@ static inline void printk_get_console_flush_type(struct console_flush_type *ft)
>   				ft->nbcon_offload = true;
>   			else
>   				ft->nbcon_atomic = true;
> +			ft->nbcon_atomic = true;
>   		}
>   
>   		/* Legacy consoles are flushed directly when possible. */
> diff --git a/kernel/printk/nbcon.c b/kernel/printk/nbcon.c
> index 4b03b019cd5ee..152d9ac28176d 100644
> --- a/kernel/printk/nbcon.c
> +++ b/kernel/printk/nbcon.c
> @@ -1185,6 +1185,7 @@ static bool nbcon_kthread_should_wakeup(struct console *con, struct nbcon_contex
>   
>   	if (kthread_should_stop())
>   		return true;
> +	return false;
>   
>   	/*
>   	 * Block the kthread when the system is in an emergency or panic mode.
> ===== END HACK =====
> 
> I am expecting that this will work and the warning will be cleanly
> visible on the UART console.

With this I see a little more output on the console, but it still
appears to hang later and never fully boots.

Boot log here: https://pastebin.com/Xe1vmeUj
  
> If the above test was successful, I am wondering if cpuidle should be
> disabling the deferred printing as this seems quite similar to suspend.
> 
> Could you perform another test using the following hack _instead_ of the
> above hack? This is probably a bit extreme since it does more than just
> disable deferring, but it could help us figure out what is going on.
> 
> ===== BEGIN HACK2 =====
> diff --git a/drivers/cpuidle/cpuidle-tegra.c b/drivers/cpuidle/cpuidle-tegra.c
> index aca907a62bb5d..0fb449e23f677 100644
> --- a/drivers/cpuidle/cpuidle-tegra.c
> +++ b/drivers/cpuidle/cpuidle-tegra.c
> @@ -15,6 +15,7 @@
>   #define pr_fmt(fmt)	"tegra-cpuidle: " fmt
>   
>   #include <linux/atomic.h>
> +#include <linux/console.h>
>   #include <linux/cpuidle.h>
>   #include <linux/cpumask.h>
>   #include <linux/cpu_pm.h>
> @@ -165,6 +166,8 @@ static __cpuidle int tegra_cpuidle_state_enter(struct cpuidle_device *dev,
>   {
>   	int err;
>   
> +	console_suspend_all();
> +
>   	/*
>   	 * CC6 state is the "CPU cluster power-off" state.  In order to
>   	 * enter this state, at first the secondary CPU cores need to be
> @@ -176,7 +179,7 @@ static __cpuidle int tegra_cpuidle_state_enter(struct cpuidle_device *dev,
>   	if (index == TEGRA_CC6) {
>   		err = tegra_cpuidle_coupled_barrier(dev);
>   		if (err)
> -			return err;
> +			goto out;
>   	}
>   
>   	local_fiq_disable();
> @@ -204,6 +207,8 @@ static __cpuidle int tegra_cpuidle_state_enter(struct cpuidle_device *dev,
>   	cpu_pm_exit();
>   	tegra_pm_clear_cpu_in_lp2();
>   	local_fiq_enable();
> +out:
> +	console_resume_all();
>   
>   	return err ?: index;
>   }
> diff --git a/kernel/printk/printk.c b/kernel/printk/printk.c
> index 3fcdf4b4e2e53..8d126b5068684 100644
> --- a/kernel/printk/printk.c
> +++ b/kernel/printk/printk.c
> @@ -2787,7 +2787,7 @@ void console_suspend_all(void)
>   	struct console *con;
>   
>   	if (console_suspend_enabled)
> -		pr_info("Suspending console(s) (use no_console_suspend to debug)\n");
> +		pr_info_once("Suspending console(s) (use no_console_suspend to debug)\n");
>   
>   	/*
>   	 * Flush any console backlog and then avoid queueing irq_work until
> ===== END HACK2 =====


I tried the above anyway, but this is hitting a 'scheduling while atomic' bug ...

[    2.083585] BUG: scheduling while atomic: swapper/0/0/0x00000002
[    2.084118] Modules linked in:
[    2.099178]
[    2.106346] CPU: 1 UID: 0 PID: 0 Comm: swapper/1 Not tainted 7.2.0-next-20260817-00003-ga0c285a6f831 #23 PREEMPT
[    2.106413] Hardware name: NVIDIA Tegra SoC (Flattened Device Tree)
[    2.106440] Call trace:
[    2.106468]  unwind_backtrace from show_stack+0x10/0x14
[    2.106596]  show_stack from dump_stack_lvl+0x54/0x68
[    2.106714]  dump_stack_lvl from __schedule_bug+0x50/0x64
[    2.106815]  __schedule_bug from __schedule+0x9e0/0x1348
[    2.106897]  __schedule from schedule+0x2c/0xf8
[    2.106962]  schedule from schedule_timeout+0xbc/0x114
[    2.107049]  schedule_timeout from __down_common+0x14c/0x3ac
[    2.107137]  __down_common from down+0x4c/0x60
[    2.107213]  down from __pr_flush+0x294/0x3d0
[    2.107284]  __pr_flush from console_suspend_all+0x44/0xc8
[    2.107342]  console_suspend_all from tegra_cpuidle_state_enter+0x18/0x33c
[    2.107406]  tegra_cpuidle_state_enter from tegra_cpuidle_enter+0x168/0x184
[    2.107470]  tegra_cpuidle_enter from cpuidle_enter_state+0x80/0x610
[    2.107533]  cpuidle_enter_state from cpuidle_enter_state_coupled+0x3ac/0x400
[    2.107637]  cpuidle_enter_state_coupled from cpuidle_enter+0x3c/0x54
[    2.107736]  cpuidle_enter from do_idle+0x150/0x340
[    2.107809]  do_idle from cpu_startup_entry+0x28/0x2c
[    2.107861]  cpu_startup_entry from secondary_start_kernel+0x118/0x120
[    2.107959]  secondary_start_kernel from 0x1160a0
[    2.108085] bad: scheduling from the idle thread!

Cheers
Jon

-- 
nvpublic
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.