Re: [PATCH 08/17] mpi3mr: Fix out-of-bounds bitmap access during device removal

Chandrakanth Patil <[email protected]>
Newsgroups org.kernel.vger.linux-scsi,dev.linux.lists.sashiko-reviews
Message-ID <CABvwm=NVp8RmHhwM4dSt1wRSLrZky0_s102t0=niRMeo6DMA5A@mail.gmail.com>
On Tue, Aug 18, 2026 at 11:43 PM <[email protected]> wrote:
>
> Thank you for your contribution! Sashiko AI review found 3 potential issue(s) to consider:
>
> New issues:
> - [High] Placement of the newly added bounds check introduces a delayed command queue stall and async chain breakage.
>
> Pre-existing issues:
> - [Critical] Concurrent lockless modification of delayed lists from process and IRQ contexts.
> - [High] TOCTOU race window during delayed command hardware dispatch allows double-free.
> --
>
> --- Patch [8]: [PATCH 08/17] mpi3mr: Fix out-of-bounds bitmap access during device removal ---
> commit 498a49c2b9e406ac25107bbe9cd216b5c786f292
> Author: Chandrakanth Patil <[email protected]>
>
> mpi3mr: Fix out-of-bounds bitmap access during device removal
>
> Device handles reported by the controller are used to index the remove
> pending bitmap and to build a task management request without being
> compared against the maximum handle the controller reported.
>
> Check the handle before using it.
>
> > diff --git a/drivers/scsi/mpi3mr/mpi3mr_os.c b/drivers/scsi/mpi3mr/mpi3mr_os.c
> > index f80a21ec161b6..7ae4b4fad3f4a 100644
> > --- a/drivers/scsi/mpi3mr/mpi3mr_os.c
> > +++ b/drivers/scsi/mpi3mr/mpi3mr_os.c
> > @@ -2401,7 +2401,8 @@ static void mpi3mr_dev_rmhs_complete_iou(struct mpi3mr_ioc *mrioc,
> >               ioc_info(mrioc,
> >                   "%s :dev removal handshake completed successfully: handle(0x%04x)\n",
> >                   __func__, drv_cmd->dev_handle);
> > -             clear_bit(drv_cmd->dev_handle, mrioc->removepend_bitmap);
> > +             if (drv_cmd->dev_handle < mrioc->facts.max_devhandle)
> > +                     clear_bit(drv_cmd->dev_handle, mrioc->removepend_bitmap);
> >       }
> >
> >       if (!list_empty(&mrioc->delayed_rmhs_list)) {
>
> [Severity: High]
> This is a pre-existing issue, but is there a race window here if the
> hardware request completes extremely quickly?
>
> mpi3mr_dev_rmhs_send_tm() sends the request to hardware. Only after it
> returns does the code list_del() and kfree() the delayed_dev_rmhs node:
>
> mpi3mr_dev_rmhs_complete_iou()
> {
>         ...
>         mpi3mr_dev_rmhs_send_tm(mrioc, drv_cmd->dev_handle, drv_cmd,
>             drv_cmd->iou_rc);
>         list_del(&delayed_dev_rmhs->list);
>         kfree(delayed_dev_rmhs);
>         return;
> }
>
> If the completion executes concurrently and triggers another call to
> mpi3mr_dev_rmhs_complete_iou(), could it dequeue and dispatch the exact
> same node before the first thread has a chance to remove it, leading to
> a double-free?
>
> Similar patterns seem to exist in mpi3mr_complete_evt_ack().
>
> [Severity: Critical]
> This is also a pre-existing issue, but I noticed there might be unprotected
> list operations on delayed_rmhs_list.
>
> The delayed_rmhs_list is appended to in mpi3mr_dev_rmhs_send_tm():
>
>         list_add_tail(&delayed_dev_rmhs->list,
>             &mrioc->delayed_rmhs_list);
>
> And it is removed from in mpi3mr_dev_rmhs_complete_iou():
>
>         list_del(&delayed_dev_rmhs->list);
>
> Are these list operations safe from concurrent modifications between process
> and IRQ contexts, or should they be protected by a spinlock to prevent
> list corruption? Similar patterns exist for delayed_evtack_cmds_list.
>
> [ ... ]
>
> > @@ -2564,6 +2565,12 @@ static void mpi3mr_dev_rmhs_send_tm(struct mpi3mr_ioc *mrioc, u16 handle,
> >       drv_cmd->callback = mpi3mr_dev_rmhs_complete_tm;
> >       drv_cmd->dev_handle = handle;
> >       drv_cmd->iou_rc = iou_rc;
> > +     if (handle >= mrioc->facts.max_devhandle) {
> > +             ioc_err(mrioc, "dev_remove_hs: handle(0x%04x) >= max_devhandle(0x%04x)\n",
> > +                     handle, mrioc->facts.max_devhandle);
> > +             goto out_failed;
> > +     }
> > +
>
> [Severity: High]
> If handle >= mrioc->facts.max_devhandle, the code goes to out_failed.
> This frees the command slot synchronously but does not invoke the
> completion callback.
>
> If this was called from mpi3mr_dev_rmhs_complete_iou() to process a delayed
> command, would this synchronous failure prevent any subsequent delayed
> items from being processed since the completion chain is broken?
>
> Also, because this check is placed here, out-of-bounds handles can still be
> appended to the delayed_rmhs_list earlier in this function when command
> slots are full:
>
> mpi3mr_dev_rmhs_send_tm()
> {
>         ...
>         if (cmd_idx >= MPI3MR_NUM_DEVRMCMD) {
>                 delayed_dev_rmhs = kzalloc_obj(*delayed_dev_rmhs, GFP_ATOMIC);
>                 ...
>                 delayed_dev_rmhs->handle = handle;
>                 list_add_tail(&delayed_dev_rmhs->list, ...);
>                 return;
>         }
>         ...
> }
>
> Should the bounds check happen before allocating and adding to the delayed
> list?
>
> >       tm_req.dev_handle = cpu_to_le16(handle);
> >       tm_req.task_type = MPI3_SCSITASKMGMT_TASKTYPE_TARGET_RESET;
> >       tm_req.host_tag = cpu_to_le16(drv_cmd->host_tag);
>
> --
> Sashiko AI review · https://sashiko.dev/#/patchset/[email protected]?part=8

Thanks for the review.

New issues:
- The observation regarding the placement of the handle bounds check
is valid and will be fixed in the v2 submission by moving the check to
the entry of mpi3mr_dev_rmhs_send_tm().

Pre-existing issues:
- The lockless list operations and race window in delayed command
queue handling are pre-existing issues in the driver and will be
addressed in a separate follow-up patch series.
smime.p7s (application/pkcs7-signature, 5.4 KB) - not displayed
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.