Re: [BUG] two raid consistency bugs

Qu Wenruo <[email protected]>
Newsgroups org.kernel.vger.linux-btrfs
Message-ID <[email protected]>

在 2026/7/15 15:16, Zhang Boyang 写道:
> Hi,
> 
> On 2026/7/15 05:40, Qu Wenruo wrote:
>>> At first power failure during transaction N, metadata trees of
>>> generation N are written to disk A, but super is not committed. Nothing
>>> is written to disk B.
>>>
>>> At second power failure during a different transaction N,
>>
>> If it's a different transaction, why it will still have the same 
>> transid N?
>>
> 
> Because the power failure occurred just before writing superblock 
> (transid N). The superblock on disk still has transid N-1.
> 
> After reboot, looking at superblock which transid is N-1, btrfs has no 
> idea of transid N existed previously, so it uses transid N for new 
> transcation.

Then there should be no problem at least at the next mount after the 
power loss.

> 
>>> nothing is
>>> written to disk A, but metadata trees and super is committed to disk B.
>>>
>>> This creates a ambiguous generation N in two disks. Currently btrfs
>>> can't detect this, and can lead to severe damages.

At the next mount, btrfs should detect device B has the latest super 
block, and use that as the super block to mount.

Since metadata are all written to device B, even disk A may have some 
stale tree blocks with transid N, stale tree blocks still need to meet 
other conditions like root owner, level, first key checks.

I won't say that's impossible, and won't say we shouldn't do anything to 
address it, but this is a variant of the split brain problems mentioned 
in the past.

I strongly recommend to find out that thread and check if any of the 
ideas are explored before and if they have their limits.

> 
> By the way, I came up another solution:
> 
> If generation mismatch between devices (or log-tree mismatch) is 
> detected at mount time, set a dirty flag in superblock on device which 
> is behind.
> 
> If dirty flag is set for a device, disable read load balancing for that 
> device. So (meta)data only read from latest device(s).

What if some metadata only arrives at that stale device, but not 
completely reached the good device?

That kills the only chance to get the good metadata mirror.

And this won't solve the split brain situation either.

> 
> If dirty flag is detected, ask user to run a scrub. The dirty flag is 
> cleared after a successful scrub.
> 
> 
> Zhang Boyang
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.