Re: [LTP] [PATCH v5] min_free_kbytes: Handle transient memory drops in check_monitor

Petr Vorel <[email protected]>
Newsgroups it.linux.lists.ltp
Message-ID <20260806103146.GA838915@pevik>
Hi Wei,

> High memory pressure can cause MemFree to temporarily drop below the
> min_free_kbytes threshold before the kernel reclaimer can catch up.
> This results in intermittent test failures, observed on openQA aarch64
> virtual machines.

> Implement a 2-second grace period with high-accuracy 10ms fixed polling
> in check_monitor() to allow the kernel time to reclaim memory.

> Introduce a 10% tolerance (90% threshold) for the MemFree check.
> Measurements under extreme pressure show MemFree can dip as low as
> ~50% to ~70% of the target. While it typically recovers above 90%
> within one second, hitting the exact 100% watermark sometimes can take
> significantly longer. This tolerance prevents false positives during the
> slow recovery tail while still ensuring memory is maintained near the
> required level.

> Also, increase the monitor's idle polling frequency from 2s to 100ms
> to improve responsiveness during the test run.

+1

> Enhanced diagnostics are added to report the minimum memory level seen
> during the pressure period to aid in future calibration.

> +++ b/testcases/kernel/mem/tunable/min_free_kbytes.c
...
>  	while (!end) {
>  		memfree = SAFE_READ_MEMINFO("MemFree:");
>  		tune = TST_SYS_CONF_LONG_GET(PATH_VM_MIN_FREE_KBYTES);
> +		/*
> +		 * Allow 10% tolerance to account for transient states.
> +		 */
> +		threshold = tune * 9 / 10;

>  		if (memfree < tune) {
> -			tst_res(TINFO, "MemFree is %lu kB, "
> -				"min_free_kbytes is %lu kB", memfree, tune);
> -			tst_res(TFAIL, "MemFree < min_free_kbytes");
> -			violated = 1;
> +			min_memfree = memfree;
> +			/*
> +			 * Give it some time to reclaim. The kernel should keep
> +			 * MemFree above min_free_kbytes, but transient drops
> +			 * are possible under high pressure.
> +			 * Check every 10ms for up to 2 seconds for high accuracy.
> +			 */
> +			for (i = 10; i <= 2000; i += 10) {
> +				usleep(10000);
You used i for usleep() in v1 and v2 [1]. Since v3 [2] you keep using
usleep(10000), therefore we could have just simple loop (readability):

			for (i = 0; i < 200; i++) {

Otherwise LGTM, thanks! (I don't agree with AI that 90% threshold is
hiding bugs). Please repost, so that we have patch in the new patchwork instance
in patchwork.kernel.org (+ please create an user on that instance if you haven't
already done), because old patchwork don't receive emails any more.

Reviewed-by: Petr Vorel <[email protected]>

Kind regards,
Petr

[1] https://lore.kernel.org/ltp/[email protected]/
[2] https://lore.kernel.org/ltp/[email protected]/

-- 
Mailing list info: https://lists.linux.it/listinfo/ltp
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.