Re: [PATCH 1/2] ext4: fix readdir position truncation on 32-bit kernels
Jan Kara <[email protected]> Thu, 6 Aug 2026 11:49:31 +0200
| Newsgroups | org.kernel.vger.linux-ext4,dev.linux.lists.ocfs2-devel,org.kernel.vger.linux-kernel,org.kernel.vger.stable |
|---|---|
| Message-ID | <jodqssdetawej6xytscjgfia72awx4exvjhdorksgiimrpfilw@nddk3kvjipn4> |
On Thu 06-08-26 10:20:43, Zhan Xusheng wrote:
> In ext4_readdir(), the directory cookie position is rebuilt with
>
> ctx->pos = (ctx->pos & ~(sb->s_blocksize - 1)) | offset;
>
> `ctx->pos` is loff_t (signed 64-bit), while `sb->s_blocksize` is
> unsigned long. On 32-bit kernels unsigned long is 32-bit, so the mask
>
> ~(sb->s_blocksize - 1)
>
> is computed as a 32-bit unsigned value (e.g. 0xfffff000 for a 4 KiB
> block size). In the AND expression with the 64-bit `ctx->pos`, that
> unsigned operand is zero-extended to 64 bits per the usual arithmetic
> conversions, yielding 0x00000000fffff000. The high 32 bits of
> `ctx->pos` are silently cleared, even though directory size is
> allowed to exceed 4 GiB on 32-bit (s_maxbytes for ext4 is many TiB).
>
> When readdir() crosses the 4 GiB boundary on a 32-bit kernel the
> position is reset back into the first 4 GiB block, making the
> re-validation path re-enumerate already-returned dirents indefinitely.
>
> ext4_readdir() reaches this linear path for non-indexed directories, and
> as the fallback after ext4_dx_readdir() returns ERR_BAD_DX_DIR, so a
> directory large enough to cross 4 GiB can hit it.
>
> This is the same class of bug that commit 3dce5bb82c97 ("exfat: Fix
> bitwise operation having different size") fixed in exfat. Cast the
> operand to loff_t so the mask is 64-bit before the AND:
>
> ctx->pos = (ctx->pos & ~((loff_t)sb->s_blocksize - 1)) | offset;
>
> 64-bit kernels are unaffected (unsigned long is 64-bit there, no
> truncation occurs).
>
> The truncation was confirmed with a freestanding 32-bit test program
> mirroring the kernel expression: input ctx->pos = 0x100000100 produces
> output 0x100 with the unfixed expression and 0x100000100 with the
> cast.
>
> Fixes: ac27a0ec112a ("[PATCH] ext4: initial copy of files from ext3")
> Cc: [email protected]
> Signed-off-by: Zhan Xusheng <[email protected]>
I'll note this is a very theoretical issue. I don't think your life it long
enough to create a 4GB non-indexed directory in ext4 :) (due to quadratic
complexity of the adding of directory entry). But the fix is right so feel
free to add:
Reviewed-by: Jan Kara <[email protected]>
Honza
> ---
> fs/ext4/dir.c | 2 +-
> 1 file changed, 1 insertion(+), 1 deletion(-)
>
> diff --git a/fs/ext4/dir.c b/fs/ext4/dir.c
> index 17edd678fa87..8113f43d4989 100644
> --- a/fs/ext4/dir.c
> +++ b/fs/ext4/dir.c
> @@ -252,7 +252,7 @@ static int ext4_readdir(struct file *file, struct dir_context *ctx)
> sb->s_blocksize);
> }
> offset = i;
> - ctx->pos = (ctx->pos & ~(sb->s_blocksize - 1))
> + ctx->pos = (ctx->pos & ~((loff_t)sb->s_blocksize - 1))
> | offset;
> info->cookie = inode_query_iversion(inode);
> }
> --
> 2.43.0
>
--
Jan Kara <[email protected]>
SUSE Labs, CR