Re: [PATCH 09/15] sched_ext: Generalize the reject DSQ reenqueue path

[email protected]
Newsgroups dev.linux.lists.sched-ext
Message-ID <[email protected]>
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] Re-rejecting a task triggers false-positive `WARN_ON_ONCE`s and permanently loses the rejection reason due to conflicting mask management.
--

--- Patch [9]: [PATCH 09/15] sched_ext: Generalize the reject DSQ reenqueue path ---
commit 0787bee4f7073d43925067d6fff35bd1a30ed061
Author: Andrea Righi <[email protected]>

sched_ext: Generalize the reject DSQ reenqueue path

This commit generalizes the reject DSQ re-enqueue path by making it
unconditional and carrying the re-enqueue reason directly in p->scx.flags.

> diff --git a/kernel/sched/ext/ext.c b/kernel/sched/ext/ext.c
> --- a/kernel/sched/ext/ext.c
> +++ b/kernel/sched/ext/ext.c

[ ... ]

> @@ -4598,23 +4600,22 @@ static void scx_reenq_reject(struct rq *rq)
>  
>  	lockdep_assert_rq_held(rq);
>  
> -	if (!scx_has_subs() || list_empty(&rq->scx.reject_dsq.list))
> +	if (list_empty(&rq->scx.reject_dsq.list))
>  		return;
>  
>  	/*
> -	 * Move to a private list so a task re-rejected by the
> +	 * Move tasks to a private list so a task re-rejected by
>  	 * scx_do_enqueue_task() below isn't revisited this round.
>  	 */
>  	list_for_each_entry_safe(p, n, &rq->scx.reject_dsq.list, scx.dsq_list.node) {
> +		u32 reason = p->scx.flags & SCX_TASK_REENQ_REASON_MASK;
> +
>  		/* migration_pending tasks should have bypassed to local DSQ */
> -		if (WARN_ON_ONCE(p->migration_pending))
> -			continue;
> +		WARN_ON_ONCE(p->migration_pending);
> +		WARN_ON_ONCE(!reason);

[Severity: High]
Will this WARN_ON_ONCE(!reason) trigger on the next tick if a task is
synchronously re-rejected during scx_do_enqueue_task()?

When a task is pushed back to the reject_dsq during the enqueue attempt,
its rejection reason mask seems to get cleared immediately after by the code
below, leaving the reason as 0 on the next drain.

>  
>  		scx_dispatch_dequeue(rq, p);
> -
> -		if (WARN_ON_ONCE(p->scx.flags & SCX_TASK_REENQ_REASON_MASK))
> -			p->scx.flags &= ~SCX_TASK_REENQ_REASON_MASK;
> -		p->scx.flags |= SCX_TASK_REENQ_CAP;
> +		p->scx.flags |= reason;
>  
>  		list_add_tail(&p->scx.dsq_list.node, &tasks);
>  	}
> @@ -4627,9 +4628,6 @@ static void scx_reenq_reject(struct rq *rq)
>  		p->scx.flags &= ~SCX_TASK_REENQ_REASON_MASK;

[Severity: High]
Does unconditionally clearing the mask here corrupt the state of newly
re-rejected tasks that were just put back on the reject_dsq?

If scx_do_enqueue_task() re-rejects the task, it gets pushed back via
scx_resolve_local_dsq().

Clearing the mask here wipes out the rejection reason while the task sits
on the reject_dsq. Without a valid rejection reason, the BPF scheduler
cannot correct the placement (e.g., missing capabilities).

Could this lead to repeated rejections until the task hits
SCX_REENQ_MAX_REPEAT, triggering __scx_exit() and crashing the BPF scheduler?

A similar issue seems to exist in reenq_local(), where the mask is also
unconditionally cleared after scx_do_enqueue_task() returns.

>  	}
>  }
> -#else
> -static void scx_reenq_reject(struct rq *rq) {}
> -#endif

[ ... ]

> diff --git a/kernel/sched/ext/sub.c b/kernel/sched/ext/sub.c
> --- a/kernel/sched/ext/sub.c
> +++ b/kernel/sched/ext/sub.c
> @@ -752,6 +752,9 @@ struct scx_dispatch_q *scx_resolve_local_dsq(struct scx_sched *sch, struct rq *r
>  
>  	p->scx.reenq_reason_caps = missing;
>  	p->scx.reenq_reason_cid = cid;
> +	if (WARN_ON_ONCE(p->scx.flags & SCX_TASK_REENQ_REASON_MASK))
> +		p->scx.flags &= ~SCX_TASK_REENQ_REASON_MASK;
> +	p->scx.flags |= SCX_TASK_REENQ_CAP;

[Severity: High]
Will this WARN_ON_ONCE fire incorrectly on valid re-rejections?

scx_reenq_reject() deliberately preserves the reason mask prior to the
enqueue attempt. If a task gets synchronously re-rejected here, it will
still have its mask set, which triggers this warning.

>  
>  	return &rq->scx.reject_dsq;
>  }

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/[email protected]?part=9
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.