Re: [PATCH v6 03/15] s390/vfio-ap: Functions to initialize/release vfio device migration data

Anthony Krowiak <[email protected]> Tue, 4 Aug 2026 15:51:44 -0400
Newsgroups dev.linux.lists.sashiko-reviews,org.kernel.vger.kvm,org.kernel.vger.linux-s390
Message-ID <[email protected]>

On 7/27/26 1:48 PM, [email protected] wrote:
> Thank you for your contribution! Sashiko AI review found 3 potential issue(s) to consider:
>
> New issues:
> - [High] The driver prematurely exposes the VFIO migration capability to userspace while only providing stub implementations that silently pretend to succeed.
>
> Pre-existing issues:
> - [High] The eventfd contexts for `req_trigger` and `cfg_chg_trigger` are leaked when the matrix mdev is removed or the device is closed.
> - [Medium] In `vfio_ap_mdev_probe()`, the device is exposed to userspace before `req_trigger` and `cfg_chg_trigger` are initialized, allowing them to be clobbered by a concurrent ioctl.
> --
>
> --- Patch [3]: [PATCH v6 03/15] s390/vfio-ap: Functions to initialize/release vfio device migration data ---
> commit faa8a565ea108d20c1807c07fbe1be6336a97435
> Author: Anthony Krowiak <[email protected]>
>
> s390/vfio-ap: Functions to initialize/release vfio device migration data
>      
> Provides the functions that initialize and release the data structures
> used during live guest migration.
>
>> diff --git a/drivers/s390/crypto/vfio_ap_migration.c b/drivers/s390/crypto/vfio_ap_migration.c
>> index fc3b2e154ad8b..ae0bd1f415feb 100644
>> --- a/drivers/s390/crypto/vfio_ap_migration.c
>> +++ b/drivers/s390/crypto/vfio_ap_migration.c
> [ ... ]
>
>> +static struct file *vfio_ap_set_state(struct vfio_device *vdev,
>> +				      enum vfio_device_mig_state  new_state)
>> +{
>> +	return NULL;
>> +}
> [Severity: High]
> Does returning NULL unconditionally here trick the VFIO FSM into believing the
> state transition succeeded immediately without actually transferring any data?

This is only a stub function here because the vfio_ap_migration_ops are set
in the vfio_device mig_ops object in the 
vfio_ap_init_migration_capabilities (below).
This function is fully implemented in patch 05/15. This patch series 
implements
in a stepwise manner to limit the functionality presented in each patch 
to make
it easier to understand and review the patches.

>
> [ ... ]
>
>> +void vfio_ap_init_migration_capabilities(struct ap_matrix_mdev *matrix_mdev)
>> +{
>> +	if (ap_is_se_guest())
>> +		return;
>> +
>> +	matrix_mdev->vdev.migration_flags = VFIO_MIGRATION_STOP_COPY;
>> +	matrix_mdev->vdev.mig_ops = &vfio_ap_migration_ops;
>> +}
> [Severity: High]
> Since the state transition operations like vfio_ap_set_state() currently just
> return NULL, does advertising the VFIO_MIGRATION_STOP_COPY capability here
> prematurely expose an incomplete migration implementation to userspace, which
> could result in broken live migration and guest data corruption?

No. This follows directly from the response above; migration is not possible
until the entire patch series is applied. Setting migration_flags to
VFIO_MIGRATION_STOP_COPY indicates that STOP,  STOP_COPY and RESUMING
are supported which will be the case going forward. As the subject line 
implies,
the purpose of this patch is limited to initializing the vfio device 
migration data.

>
>> diff --git a/drivers/s390/crypto/vfio_ap_ops.c b/drivers/s390/crypto/vfio_ap_ops.c
>> index 36786d70a88f5..90b0fce0123b5 100644
>> --- a/drivers/s390/crypto/vfio_ap_ops.c
>> +++ b/drivers/s390/crypto/vfio_ap_ops.c
> [ ... ]
>
>> +static void vfio_ap_mdev_release_dev(struct vfio_device *vdev)
>> +{
>> +	struct ap_matrix_mdev *matrix_mdev;
>> +
>> +	mutex_lock(&matrix_dev->mdevs_lock);
>> +	matrix_mdev = container_of(vdev, struct ap_matrix_mdev, vdev);
>> +	vfio_ap_release_migration_data(matrix_mdev);
>> +	mutex_unlock(&matrix_dev->mdevs_lock);
>> +}
> [Severity: High]
> This isn't a bug introduced by this patch, but does vfio_ap_mdev_release_dev()
> fail to put the req_trigger and cfg_chg_trigger eventfd contexts?
>
> If they are also not cleaned up in vfio_ap_mdev_close_device(), could this
> allow a user to systematically leak kernel memory by repeatedly opening the
> device, configuring IRQs, and closing it?

This is a legitimate observation. It will be fixed with an individual 
fix patch
unrelated to this series.

>
> [ ... ]
>
>>   	ret = vfio_register_emulated_iommu_dev(&matrix_mdev->vdev);
>>   	if (ret)
>>   		goto err_put_vdev;
>> +
>> +	mutex_lock(&matrix_dev->mdevs_lock);
>>   	matrix_mdev->req_trigger = NULL;
>>   	matrix_mdev->cfg_chg_trigger = NULL;
> [Severity: Medium]
> This isn't a bug introduced by this patch, but does assigning NULL to
> req_trigger and cfg_chg_trigger here risk clobbering valid pointers?
>
> If a concurrent ioctl configures the triggers via VFIO_DEVICE_SET_IRQS
> immediately after the device is exposed via vfio_register_emulated_iommu_dev(),
> but before we acquire mdevs_lock here, could we overwrite the initialized
> triggers and leak the eventfd contexts?

The trigger assignments in vfio_ap_set_request_irq and
vfio_ap_set_cfg_change_irq are only reachable via vfio_ap_mdev_ioctl,
which is only callable after open_device succeeds. open_device calls
vfio_ap_init_migration_data which requires the device to be registered
and opened — that open hasn't occurred yet at probe time. The window
between vfio_register_emulated_iommu_dev and the NULL assignments exists
but no ioctl path can reach req_trigger during it. The NULL assignments
are also redundant since vfio_alloc_device (line 759) zero-initialises
the whole matrix_mdev struct, but they are harmless.

>
>>   	dev_set_drvdata(&mdev->dev, matrix_mdev);
>> -	mutex_lock(&matrix_dev->mdevs_lock);
>>   	list_add(&matrix_mdev->node, &matrix_dev->mdev_list);
>>   	mutex_unlock(&matrix_dev->mdevs_lock);
>>   	return 0;