Re: [PATCH v5 4/6] vhost-user-blk: make inflight-migration prop mutable

Alexandr Moshkov <[email protected]> Thu, 30 Jul 2026 14:01:02 +0500
Newsgroups dev.linux.lists.virtio-fs,org.nongnu.qemu-devel
Message-ID <[email protected]>
On 7/29/26 23:50, Michael S. Tsirkin wrote:
> On Wed, Jul 29, 2026 at 08:18:19PM +0500, Alexandr Moshkov wrote:
>> On 7/29/26 19:48, Michael S. Tsirkin wrote:
>>> On Wed, Jul 29, 2026 at 07:07:11PM +0500, Alexandr Moshkov wrote:
>>>> On 7/29/26 17:13, Michael S. Tsirkin wrote:
>>>>> On Wed, Jul 29, 2026 at 03:56:54PM +0500, Alexandr Moshkov wrote:
>>>>>> On 7/29/26 15:48, Michael S. Tsirkin wrote:
>>>>>>
>>>>>>        On Wed, Jul 29, 2026 at 03:35:31PM +0500, Alexandr Moshkov wrote:
>>>>>>
>>>>>>            On 7/29/26 15:14, Michael S. Tsirkin wrote:
>>>>>>
>>>>>>                On Wed, Jul 29, 2026 at 03:09:57PM +0500, Alexandr Moshkov wrote:
>>>>>>
>>>>>>                    On 7/29/26 14:32, Michael S. Tsirkin wrote:
>>>>>>
>>>>>>                         On Tue, Jul 28, 2026 at 03:08:39PM +0500, Alexandr Moshkov wrote:
>>>>>>
>>>>>>                             When migrating from a QEMU version that supports inflight-migration to
>>>>>>                             an older one that does not, there is no way to disable the feature at
>>>>>>                             runtime — the VM must be stopped and reconfigured. This is impractical
>>>>>>                             in production environments.
>>>>>>
>>>>>>                             Make the inflight-migration property mutable after device realization
>>>>>>                             so it can be toggled via qom-set without restarting the VM.
>>>>>>
>>>>>>                             Acked-by: Raphael Norwitz <[email protected]>
>>>>>>                             Signed-off-by: Alexandr Moshkov <[email protected]>
>>>>>>
>>>>>>
>>>>>>
>>>>>>                         I think I am beginning to understand.
>>>>>>
>>>>>>
>>>>>>                         You are running qemu with inflight-migration on and want to migrate
>>>>>>                         to qemu without inflight-migration at all.
>>>>>>
>>>>>>
>>>>>>                         Since it is guest transparent you could retrofit it like this.
>>>>>>
>>>>>>                         My question is why is it worth it, we do not normally support
>>>>>>                         migrating between qemu versions with different command lines.
>>>>>>
>>>>>>                    I think you right. Maybe I've been focusing too much on the ability to migrate
>>>>>>                    between versions of qemu with and without inflight migration.
>>>>>>
>>>>>>                    This series allows to turn off and on the inflight-migration feature at runtime
>>>>>>                    without recreating the VM
>>>>> restarting you mean.
>>>> Yes, sorry.
>>>>>> (it was the only way to turn off the feature, since
>>>>>>                    the protocol feature could no longer be turned off, after initialization with
>>>>>>                    the backend). This is more important feature, and it leads to the fact that
>>>>>>                    this feature allows to migrate between different versions of qemu (with and
>>>>>>                    without inflight-migration support).
>>>>> so again it's not the 1st feature we have like this. we tie migration
>>>>> to the machine type and same set of command line flags
>>>>> specifically to keep things manageable.
>>>>>
>>>>> really cross version migration is a pain as it is.
>>>> I understand that, thanks! Maybe we should focus on the fact that this
>>>> series allows to turn inflight_migration on and off (without cross version
>>>> migration) ?
>>>>
>>>> Firstly, I wanted to implement this through migration capabilities, but in
>>>> one of the threads,
>>>> it was agreed that a separate capability only for vhost-user-blk is not a
>>>> good idea.
>>> I mean it's a generic thing.
>>>
>>>> Therefore, I implemented it through a parameter that can be changed using
>>>> qom-set.
>>>>
>>>>>>                             ---
>>>>>>                              hw/block/vhost-user-blk.c | 10 ++++++++--
>>>>>>                              1 file changed, 8 insertions(+), 2 deletions(-)
>>>>>>
>>>>>>                             diff --git a/hw/block/vhost-user-blk.c b/hw/block/vhost-user-blk.c
>>>>>>                             index e3b873af7c..650c004bde 100644
>>>>>>                             --- a/hw/block/vhost-user-blk.c
>>>>>>                             +++ b/hw/block/vhost-user-blk.c
>>>>>>                             @@ -619,6 +619,8 @@ static const VMStateDescription vmstate_vhost_user_blk = {
>>>>>>                                  }
>>>>>>                              };
>>>>>>
>>>>>>                             +static PropertyInfo vhost_user_blk_inflight_migration_prop;
>>>>>>                             +
>>>>>>                              static const Property vhost_user_blk_properties[] = {
>>>>>>                                  DEFINE_PROP_CHR("chardev", VHostUserBlk, chardev),
>>>>>>                                  DEFINE_PROP_UINT16("num-queues", VHostUserBlk, num_queues,
>>>>>>                             @@ -632,8 +634,9 @@ static const Property vhost_user_blk_properties[] = {
>>>>>>                                                    VIRTIO_BLK_F_WRITE_ZEROES, true),
>>>>>>                                  DEFINE_PROP_BOOL("skip-get-vring-base-on-force-shutdown", VHostUserBlk,
>>>>>>                                                   skip_get_vring_base_on_force_shutdown, false),
>>>>>>                             -    DEFINE_PROP_BOOL("inflight-migration", VHostUserBlk,
>>>>>>                             -                     inflight_migration, false),
>>>>>>                             +    DEFINE_PROP("inflight-migration", VHostUserBlk, inflight_migration,
>>>>>>                             +                vhost_user_blk_inflight_migration_prop, bool,
>>>>>>                             +                .set_default = true, .defval.u = false),
>>>>>>                              };
>>>>>>
>>>>>>                              static void vhost_user_blk_class_init(ObjectClass *klass, const void *data)
>>>>>>                             @@ -665,6 +668,9 @@ static const TypeInfo vhost_user_blk_info = {
>>>>>>
>>>>>>                              static void virtio_register_types(void)
>>>>>>                              {
>>>>>>                             +    vhost_user_blk_inflight_migration_prop = qdev_prop_bool;
>>>>>>                             +    vhost_user_blk_inflight_migration_prop.realized_set_allowed = true;
>>>>>>                             +
>>>>>>                                  type_register_static(&vhost_user_blk_info);
>>>>>>                              }
>>>>>>
>>>>>>                         So then,  for example, let us say I paused the VM, then set the flag,
>>>>>>                         now inflight is on but GET_BASE did not drain it?
>>>>>>
>>>>>>                    If I understood the question correctly, this is valid behavior. Before
>>>>>>                    migration QEMU check protocol features to understand does the backend support
>>>>>>                    inflight migration. If it does, after that QEMU migrate inflight buffer to
>>>>>>                    other VM. If it's not, return error before migration started.
>>>>>>
>>>>>>                I apologise i reverted the logic in the question.
>>>>>>
>>>>>>
>>>>>>                I start vm and inflight is on.
>>>>>>                I stop vm get base does not drain.
>>>>>>
>>>>>>            In case of VM stop, backend wait to drain all requests. Ability to perform
>>>>>>            drain or not available only during migration by skip_drain variable in
>>>>>>            vhost_user_blk_stop().
>>>>>>
>>>>>>        What if it was migration but it failed?
>>>>>>
>>>>>> If migration failed, source vm just keep using inflight region, backend
>>>>>> continue to execute inflight requests that was tried to migrate.
>>>>> Sorry, to be more clear. I mean migration succeeded but destination failed to
>>>>> start, so now vm is stopped and we are now trying migrating to a different destination.
>>>> Is this a real case? I was thinking in case of any destination errors
>>>> (migration process or starting), source vm has to work...
>>> It does not start automatically in all cases, no.
>>>
>>>> I guess in that case, inflight requests can only be executed on different
>>>> destination or in current destination if vm be able to start.
>>>>
>>>>> IIUC currently on VM stop we simply send GET_BASE and this stops
>>>>> backend, and depending on features things remain in the inflight buffer.
>>>>> Correct me if  I am wrong.
>>>> Yes backend has to disconnect from QEMU.. So it not be changing guest
>>>> memory.
>>>>> I think that we do not want a slow flush on vm stop if we
>>>>> can avoid it, and we also do not want backend to keep changing
>>>>> guest memory when VM is stopped. No?
>>>> Yes, it is
>>>>>>            skip_drain true only if inflight_migration is on and runstate is
>>>>>>            FINISH_MIGRATION.
>>>>> FINISH_MIGRATE? actually i do not see where it affects it.
>>>> Sorry, I didn't understand) Who affects who?
>>>>>>        This btw I don't much like, a stopped VM would preferably
>>>>>>        behave the same whatever the reason to stop.
>>>>>>
>>>>>> Well, I don't see any other way to implement this.
>>>>> I am not 100% sure it's implementable as described, but I guess
>>>>> you could block changing the property if VM and thus the backend
>>>>> is not connected and running.
>>>>>
>>>>>
>>>>> Which is even more complexity but at least it is consistent.
>>>> Oh, I understand what you thinking about.
>>>>
>>>> I agreed, but why implementation (in this series) is not consistent? I was
>>>> thinking that current (in current qemu version) impl is not, because
>>>> GET_VRING_BASE has two behaviors, depending on the presence of the protocol
>>>> feature (drain or not drain requests).
>>>>
>>>> Now, I think it would be more clear: one message for drain GET_VRING_BASE
>>>> (as it was before), one for not drain GET_VRING_BASE_SKIP_DRAIN.
>>> Well the message is consistent. But its use does not seem to be.
>>>
>>>>>
>>>>>
>>>>>
>>>>>
>>>>>> I think backend must perform
>>>>>> drain in all cases except live migration with all the necessary checks
>>>>>> (protocol features, special message GET_VRING_BASE_SKIP_DRAIN).
>>>>> This is not what is going on now, right? Whether it drains depends
>>>>> on protocol features not on VM state and I think we should
>>>>> keep it like this.
>>>> Well yes, now in current qemu version drain or not drain depends on protocol
>>>> feature only, that try to be negotiated only if inflight_migration is set
>>>> on.
>>>>
>>>> In patch: in case live migration (if protocol feature and  device param)
>>>> sent message GET_VRING_BASE_SKIP_DRAIN to skip drain, in other cases sent
>>>> other message - GET_VRING_BASE with draining all requests.
>>>> So drain or not drain depends on message that QEMU send to backend.
>>>>
>>>>>>                i turn inflight off.
>>>>>>                now it looks like it will happily migrate?
>>>>>>
>>>>>>            So in this case, when vm migrated with inflight off, it will leads to using
>>>>>>            GET_VRING_BASE message, that wait all requests to be drained.
>>>>>>
>>>>>>
>>>>>>                             --
>>>>>>                             2.34.1
>>>>>>
>>>>>>
>>>>> i mean it migrated with on. vm was stopped then it flipped to off.
>>>> After stop vm will send GET_VRING_BASE to backend, that lead to drain all
>>>> requests. Non of the inflight requests will be missed here, If that's what
>>>> you mean.
>>> this:
>>>
>>> +    bool skip_drain = vhost_user_blk_inflight_needed(s) &&
>>> +                      runstate_check(RUN_STATE_FINISH_MIGRATE);
>>>
>>> is what I dislike.
>>>
>>> I dislike poking at migration state. It's painful enough that we
>>> need to worry about VM running or not.
>> I understand that. But I don't know how to do it another way. We can't just
>> erase runstate check - it will possibly leads to missing inflight requests
>> during VM stop.
>>
>> And don't know how to make stop more consistent. Do you have any thoughts?
>
> Block disabling this when VM is not running.

Alright, thanks. Now I understand)

>
>>>>>
>>>>>
>>>>>