Re: [i-g-t] tests/intel/xe_exec_mix_modes: Add multi-queue LR spinner switch subtest

"Randhawa, Jagmeet" <[email protected]>
Newsgroups org.freedesktop.lists.igt-dev
Message-ID <[email protected]>
On 8/18/2026 6:50 PM, Niranjana Vishwanathapura wrote:
> On Mon, Aug 17, 2026 at 08:00:28PM +0000, Jagmeet Randhawa wrote:
>> Add exec-multi-queue-spinner-interrupted-lr, which runs preemptible
>> spinners on every queue in a multi-queue exec queue group in long
>> running mode, then submits a dma-fence job on the same engine to force
>> the group from FAULT into DMA_FENCE mode.
>>
>> Each spinner is armed with a queue-switch semaphore
>> (multi_queue_switch_on_wait + xe_spin_preempt_wait) so it yields the
>> shared engine to its sibling once scheduled, instead of busy-looping
>> and starving the rest of the group. This lets every queue in the group
>> be genuinely resident when the interrupting job forces the mode
>> switch, exercising suspend/resume of the secondary queues rather than
>> just the primary.
>>
>> The dma-fence and LR (preempt) multi-queue suspend/resume paths are
>> already covered by the userptr-invalidation tests; this adds the
>> missing FAULT-mode coverage.
>>
>> Signed-off-by: Jagmeet Randhawa <[email protected]>
>> ---
>> tests/intel/xe_exec_mix_modes.c | 111 ++++++++++++++++++++++++++++++++
>> 1 file changed, 111 insertions(+)
>>
>> diff --git a/tests/intel/xe_exec_mix_modes.c 
>> b/tests/intel/xe_exec_mix_modes.c
>> index 209825f58..d9ebb8283 100644
>> --- a/tests/intel/xe_exec_mix_modes.c
>> +++ b/tests/intel/xe_exec_mix_modes.c
>> @@ -27,6 +27,7 @@
>>
>> #define FLAG_EXEC_MODE_LR    (0x1 << 0)
>> #define FLAG_JOB_TYPE_SIMPLE    (0x1 << 1)
>> +#define FLAG_MULTI_QUEUE    (0x1 << 2)
>>
>> #define NUM_INTERRUPTING_JOBS    1
>> #define USER_FENCE_VALUE    0xdeadbeefdeadbeefull
>> @@ -34,6 +35,7 @@
>> #define SPIN_DATA        1
>> #define EXEC_DATA        2
>> #define DATA_COUNT        3
>> +#define N_GROUP_QUEUES        2
>>
>> struct data {
>>     struct xe_spin spin;
>> @@ -221,6 +223,104 @@ run_job(int fd, struct 
>> drm_xe_engine_class_instance *hwe,
>>     xe_vm_destroy(fd, vm);
>> }
>>
>> +static void
>> +run_job_multi_queue(int fd, struct drm_xe_engine_class_instance *hwe)
>> +{
>> +    struct drm_xe_sync sync = {
>> +        .flags = DRM_XE_SYNC_FLAG_SIGNAL,
>> +        .type = DRM_XE_SYNC_TYPE_USER_FENCE,
>> +        .timeline_value = USER_FENCE_VALUE,
>> +    };
>> +    struct drm_xe_exec exec = {
>> +        .num_batch_buffer = 1,
>> +        .num_syncs = 1,
>> +        .syncs = to_user_pointer(&sync),
>> +    };
>> +    uint32_t exec_queues[N_GROUP_QUEUES];
>> +    struct xe_spin *spin[N_GROUP_QUEUES];
>> +    uint64_t addr[N_GROUP_QUEUES];
>> +    uint64_t base_addr = 0x1a0000;
>> +    int64_t fence_timeout = NSEC_PER_SEC;
>> +    int64_t timeout_short = 1;
>> +    uint64_t vm_sync = 0;
>> +    size_t bo_size, spin_size;
>> +    uint32_t vm, bo;
>> +    void *map;
>> +    int i;
>> +
>> +    igt_require(xe_has_multi_queue_engine(fd));
>> +
>> +    vm = xe_vm_create(fd, DRM_XE_VM_CREATE_FLAG_LR_MODE |
>> +                  DRM_XE_VM_CREATE_FLAG_FAULT_MODE, 0);
>> +
>> +    /* exec_queues[0] is the primary, the rest join its group. */
>> +    for (i = 0; i < N_GROUP_QUEUES; i++) {
>> +        struct drm_xe_ext_set_property multi_queue = {
>> +            .base.name = DRM_XE_EXEC_QUEUE_EXTENSION_SET_PROPERTY,
>> +            .property = DRM_XE_EXEC_QUEUE_SET_PROPERTY_MULTI_GROUP,
>> +        };
>> +        uint64_t ext = to_user_pointer(&multi_queue);
>> +
>> +        multi_queue.value = i ? exec_queues[0] : 
>> DRM_XE_MULTI_GROUP_CREATE;
>> +        exec_queues[i] = xe_exec_queue_create(fd, vm, hwe, ext);
>> +    }
>> +
>> +    spin_size = xe_bb_size(fd, sizeof(struct xe_spin));
>> +    bo_size = spin_size * N_GROUP_QUEUES;
>> +    bo = xe_bo_create(fd, vm, bo_size, vram_if_possible(fd, 
>> hwe->gt_id),
>> +              DRM_XE_GEM_CREATE_FLAG_NEEDS_VISIBLE_VRAM);
>> +    map = xe_bo_map(fd, bo, bo_size);
>> +    for (i = 0; i < N_GROUP_QUEUES; i++) {
>> +        spin[i] = (struct xe_spin *)((char *)map + i * spin_size);
>> +        addr[i] = base_addr + i * spin_size;
>> +    }
>> +
>> +    sync.addr = to_user_pointer(&vm_sync);
>> +    xe_vm_bind_async(fd, vm, 0, bo, 0, base_addr, bo_size, &sync, 1);
>> +    xe_wait_ufence(fd, &vm_sync, USER_FENCE_VALUE, 0, fence_timeout);
>> +    vm_sync = 0;
>> +
>> +    for (i = 0; i < N_GROUP_QUEUES; i++) {
>> +        /* Yield the shared engine at a switch point instead of 
>> busy-spinning. */
>> +        xe_spin_init_opts(spin[i], .addr = addr[i], .preempt = true,
>> +                  .multi_queue_switch_on_wait = true);
>
> I think here 'multi_queue_switch = true' is enough probably (as we 
> have 'preempt = true')
> instead of 'multi_queue_switch_on_wait = true'. With that, we also 
> don't need the
> xe_spin_preempt_wait/nowait() calls below as well. Can you check?
You're right. Switched to .multi_queue_switch = true, since it's an 
unconditional switch point, each spinner yields to its sibling on its 
own, so the xe_spin_preempt_wait()/xe_spin_preempt_nowait() calls are no 
longer needed and are removed. Verified on hardware (NVL-P): the subtest 
passes. Done in v2.
>
>> +        xe_spin_preempt_wait(spin[i]);
>> +        sync.addr = addr[i] + (char *)&spin[i]->exec_sync - (char 
>> *)spin[i];
>> +        exec.exec_queue_id = exec_queues[i];
>> +        exec.address = addr[i];
>> +        xe_exec(fd, &exec);
>> +        xe_spin_wait_started(spin[i]);
>> +    }
>> +
>> +    /* Force the group from FAULT into DMA_FENCE mode. */
>> +    run_job(fd, hwe, EXEC_MODE_DMA_FENCE, SIMPLE_BATCH_STORE, false, 
>> NULL);
>> +
>> +    /* Spinners must have been suspended, not completed, during the 
>> switch. */
>> +    for (i = 0; i < N_GROUP_QUEUES; i++)
>> +        igt_assert_neq(0, __xe_wait_ufence(fd, &spin[i]->exec_sync,
>> +                           USER_FENCE_VALUE, exec_queues[i],
>> +                           &timeout_short));
>> +
>
> This check doesn't mean much as we are ending the spinner below. ie., 
> job is known
> to be not complete as we have not ended the spinner. But I see you 
> want to keep it
> similar to other existing tests. Perhaps we can remove it as are sure 
> that spinner
> job won't be complete unless we end it.
|Agreed, it doesn't add value since the spinner can't complete until we 
end it. Removed the check (and the now-unused timeout_short) in v2.|
>
>> +    for (i = 0; i < N_GROUP_QUEUES; i++) {
>> +        xe_spin_end(spin[i]);
>> +        xe_spin_preempt_nowait(spin[i]);
>> +    }
>> +
>> +    for (i = 0; i < N_GROUP_QUEUES; i++)
>> +        xe_wait_ufence(fd, &spin[i]->exec_sync, USER_FENCE_VALUE,
>> +                   exec_queues[i], fence_timeout);
>> +
>> +    sync.addr = to_user_pointer(&vm_sync);
>> +    xe_vm_unbind_async(fd, vm, 0, 0, base_addr, bo_size, &sync, 1);
>> +    xe_wait_ufence(fd, &vm_sync, USER_FENCE_VALUE, 0, fence_timeout);
>> +
>> +    for (i = 0; i < N_GROUP_QUEUES; i++)
>> +        xe_exec_queue_destroy(fd, exec_queues[i]);
>> +    munmap(map, bo_size);
>> +    gem_close(fd, bo);
>> +    xe_vm_destroy(fd, vm);
>> +}
>> +
>> /**
>>  * SUBTEST: exec-simple-batch-store-lr
>>  * Description: Execute a simple batch store job in long running mode
>> @@ -235,6 +335,11 @@ run_job(int fd, struct 
>> drm_xe_engine_class_instance *hwe,
>>  * SUBTEST: exec-spinner-interrupted-dma-fence
>>  * Description: Spin in dma fence mode then get interrupted by a simple
>>  *              batch store job in long running mode
>> + *
>> + * SUBTEST: exec-multi-queue-spinner-interrupted-lr
>> + * Description: Run preemptible spinners on a multi-queue exec queue 
>> group
>> + *              in long running mode, then get interrupted by a simple
>> + *              batch store job in dma fence mode
>>  */
>> static void
>> test_exec(int fd, struct drm_xe_engine_class_instance *hwe,
>> @@ -243,6 +348,11 @@ test_exec(int fd, struct 
>> drm_xe_engine_class_instance *hwe,
>>     enum engine_execution_mode engine_execution_mode;
>>     enum job_type job_type;
>>
>> +    if (flags & FLAG_MULTI_QUEUE) {
>> +        run_job_multi_queue(fd, hwe);
>> +        return;
>> +    }
>> +
>
> Actually, we should call
> igt_require(xe_engine_class_supports_multi_queue(fd, hwe->engine_class));
> here instead of calling 'igt_require(xe_has_multi_queue_engine(fd));' in
> run_job_multi_queue. Earlier in the call chain we do it, it is better.
>
> Also, add an assert here that FLAG_EXEC_MODE_LR is set as that is the
> only test MULTI_QUEUE case supports currently.
>
> Niranjana

Done. Moved the capability check into test_exec() using 
xe_engine_class_supports_multi_queue(fd, hwe->engine_class) (dropped the 
xe_has_multi_queue_engine() call in run_job_multi_queue), and added 
igt_assert(flags & FLAG_EXEC_MODE_LR) in v2.

Thanks for the review, all three addressed in v2, will send it up shortly.

Jagmeet


>
>>     if (flags & FLAG_EXEC_MODE_LR)
>>         engine_execution_mode = EXEC_MODE_LR;
>>     else
>> @@ -267,6 +377,7 @@ int igt_main()
>>         { "simple-batch-store-dma-fence", FLAG_JOB_TYPE_SIMPLE },
>>         { "spinner-interrupted-lr", FLAG_EXEC_MODE_LR },
>>         { "spinner-interrupted-dma-fence", 0 },
>> +        { "multi-queue-spinner-interrupted-lr", FLAG_EXEC_MODE_LR | 
>> FLAG_MULTI_QUEUE },
>>         { NULL },
>>     };
>>     int fd;
>> -- 
>> 2.43.0
>>
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.