Re: [PATCH v1] tty: n_tty: use kvzalloc/kvfree for line discipline data

Greg KH <[email protected]>
Newsgroups gmane.linux.serial,gmane.linux.kernel
Message-ID <2026081706-tricky-slicing-166f@gregkh>
On Mon, Aug 17, 2026 at 09:55:26PM +0800, Xin Chen wrote:
> BT enable fails intermittently with -ETIMEDOUT (-110).  The kernel log
> shows the HCI Read Local Version command was sent and the firmware
> replied with status 0x00 (logged by hci_req_cmd_complete() BT_DBG),
> but the waiter in __hci_cmd_sync_sk() never woke up and timed out
> after 10 s:
> 
>   bluetooth hci0: Opcode 0xfc00              // __hci_cmd_sync_sk
>   bluetooth hci0: opcode 0xfc00 plen 1       // hci_cmd_sync_add
>   bluetooth hci0: skb len 4                  // hci_cmd_sync_alloc
>   bluetooth hci0: length 1                   // hci_req_sync_run
>   Bluetooth: hci0 cmd_cnt 1 cmd queued 1     // hci_cmd_work
>   Bluetooth: hci0 type 1 len 4               // hci_send_frame
>   Bluetooth: opcode 0xfc00 status 0x00       // hci_req_cmd_complete
>   <-- req_skb NULL: req_complete_skb not set,
>       hci_cmd_sync_complete() never called,
>       req_status stays HCI_REQ_PEND            -->
>   <-- 10 s later: wait_event_interruptible_timeout expires -->
>   bluetooth hci0: end: err -110              // __hci_cmd_sync_sk
> 
> The root cause is that hci_send_cmd_sync() clones the sent command
> into hdev->req_skb so that hci_req_cmd_complete() can locate the
> registered completion callback.  Under memory pressure this
> skb_clone() fails, leaving hdev->req_skb NULL.  The firmware reply
> is received and processed, but hci_req_cmd_complete() finds NULL
> req_skb, so hci_cmd_sync_complete() is never called, req_status
> stays HCI_REQ_PEND, and the waiter times out with -ETIMEDOUT.
> 
> The memory pressure is caused by n_tty_open().  When a BT UART
> transport is opened, serdev_device_open() may be called multiple
> times in quick succession, each triggering n_tty_open().  n_tty_open()
> uses vzalloc() for the ~10 KB n_tty_data structure, which always
> allocates page-by-page from the buddy order-0 free list.  Repeated
> vzalloc() calls drain enough order-0 pages that the subsequent
> skb_clone(GFP_KERNEL) in hci_send_cmd_sync() cannot get a page.

So you run out of memory?  That feels wrong.

Why not just use a specific slab for this one structure if it is so
important that it never run out?  Why was this using vzalloc() in the
first place if it could fail?

And if it does fail, doesn't everything work properly, you just need to
handle that failure in userspace correctly, right?  What is failing that
you can not recover?  If we are running out of memory here for such a
tiny allocation, odds are other things are going to go wrong so
userspace better handle that.

thanks,

greg k-h
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.