Re: [PATCH v6 03/15] s390/vfio-ap: Functions to initialize/release vfio device migration data
Anthony Krowiak <[email protected]> Tue, 4 Aug 2026 15:51:44 -0400
| Newsgroups | dev.linux.lists.sashiko-reviews,org.kernel.vger.kvm,org.kernel.vger.linux-s390 |
|---|---|
| Message-ID | <[email protected]> |
On 7/27/26 1:48 PM, [email protected] wrote: > Thank you for your contribution! Sashiko AI review found 3 potential issue(s) to consider: > > New issues: > - [High] The driver prematurely exposes the VFIO migration capability to userspace while only providing stub implementations that silently pretend to succeed. > > Pre-existing issues: > - [High] The eventfd contexts for `req_trigger` and `cfg_chg_trigger` are leaked when the matrix mdev is removed or the device is closed. > - [Medium] In `vfio_ap_mdev_probe()`, the device is exposed to userspace before `req_trigger` and `cfg_chg_trigger` are initialized, allowing them to be clobbered by a concurrent ioctl. > -- > > --- Patch [3]: [PATCH v6 03/15] s390/vfio-ap: Functions to initialize/release vfio device migration data --- > commit faa8a565ea108d20c1807c07fbe1be6336a97435 > Author: Anthony Krowiak <[email protected]> > > s390/vfio-ap: Functions to initialize/release vfio device migration data > > Provides the functions that initialize and release the data structures > used during live guest migration. > >> diff --git a/drivers/s390/crypto/vfio_ap_migration.c b/drivers/s390/crypto/vfio_ap_migration.c >> index fc3b2e154ad8b..ae0bd1f415feb 100644 >> --- a/drivers/s390/crypto/vfio_ap_migration.c >> +++ b/drivers/s390/crypto/vfio_ap_migration.c > [ ... ] > >> +static struct file *vfio_ap_set_state(struct vfio_device *vdev, >> + enum vfio_device_mig_state new_state) >> +{ >> + return NULL; >> +} > [Severity: High] > Does returning NULL unconditionally here trick the VFIO FSM into believing the > state transition succeeded immediately without actually transferring any data? This is only a stub function here because the vfio_ap_migration_ops are set in the vfio_device mig_ops object in the vfio_ap_init_migration_capabilities (below). This function is fully implemented in patch 05/15. This patch series implements in a stepwise manner to limit the functionality presented in each patch to make it easier to understand and review the patches. > > [ ... ] > >> +void vfio_ap_init_migration_capabilities(struct ap_matrix_mdev *matrix_mdev) >> +{ >> + if (ap_is_se_guest()) >> + return; >> + >> + matrix_mdev->vdev.migration_flags = VFIO_MIGRATION_STOP_COPY; >> + matrix_mdev->vdev.mig_ops = &vfio_ap_migration_ops; >> +} > [Severity: High] > Since the state transition operations like vfio_ap_set_state() currently just > return NULL, does advertising the VFIO_MIGRATION_STOP_COPY capability here > prematurely expose an incomplete migration implementation to userspace, which > could result in broken live migration and guest data corruption? No. This follows directly from the response above; migration is not possible until the entire patch series is applied. Setting migration_flags to VFIO_MIGRATION_STOP_COPY indicates that STOP, STOP_COPY and RESUMING are supported which will be the case going forward. As the subject line implies, the purpose of this patch is limited to initializing the vfio device migration data. > >> diff --git a/drivers/s390/crypto/vfio_ap_ops.c b/drivers/s390/crypto/vfio_ap_ops.c >> index 36786d70a88f5..90b0fce0123b5 100644 >> --- a/drivers/s390/crypto/vfio_ap_ops.c >> +++ b/drivers/s390/crypto/vfio_ap_ops.c > [ ... ] > >> +static void vfio_ap_mdev_release_dev(struct vfio_device *vdev) >> +{ >> + struct ap_matrix_mdev *matrix_mdev; >> + >> + mutex_lock(&matrix_dev->mdevs_lock); >> + matrix_mdev = container_of(vdev, struct ap_matrix_mdev, vdev); >> + vfio_ap_release_migration_data(matrix_mdev); >> + mutex_unlock(&matrix_dev->mdevs_lock); >> +} > [Severity: High] > This isn't a bug introduced by this patch, but does vfio_ap_mdev_release_dev() > fail to put the req_trigger and cfg_chg_trigger eventfd contexts? > > If they are also not cleaned up in vfio_ap_mdev_close_device(), could this > allow a user to systematically leak kernel memory by repeatedly opening the > device, configuring IRQs, and closing it? This is a legitimate observation. It will be fixed with an individual fix patch unrelated to this series. > > [ ... ] > >> ret = vfio_register_emulated_iommu_dev(&matrix_mdev->vdev); >> if (ret) >> goto err_put_vdev; >> + >> + mutex_lock(&matrix_dev->mdevs_lock); >> matrix_mdev->req_trigger = NULL; >> matrix_mdev->cfg_chg_trigger = NULL; > [Severity: Medium] > This isn't a bug introduced by this patch, but does assigning NULL to > req_trigger and cfg_chg_trigger here risk clobbering valid pointers? > > If a concurrent ioctl configures the triggers via VFIO_DEVICE_SET_IRQS > immediately after the device is exposed via vfio_register_emulated_iommu_dev(), > but before we acquire mdevs_lock here, could we overwrite the initialized > triggers and leak the eventfd contexts? The trigger assignments in vfio_ap_set_request_irq and vfio_ap_set_cfg_change_irq are only reachable via vfio_ap_mdev_ioctl, which is only callable after open_device succeeds. open_device calls vfio_ap_init_migration_data which requires the device to be registered and opened — that open hasn't occurred yet at probe time. The window between vfio_register_emulated_iommu_dev and the NULL assignments exists but no ioctl path can reach req_trigger during it. The NULL assignments are also redundant since vfio_alloc_device (line 759) zero-initialises the whole matrix_mdev struct, but they are harmless. > >> dev_set_drvdata(&mdev->dev, matrix_mdev); >> - mutex_lock(&matrix_dev->mdevs_lock); >> list_add(&matrix_mdev->node, &matrix_dev->mdev_list); >> mutex_unlock(&matrix_dev->mdevs_lock); >> return 0;