Re: thin pool powerfail tests and data loss
Ming Hung Tsai <[email protected]> Wed, 18 Sep 2024 18:20:45 +0800
| Newsgroups | gmane.linux.lvm.devel |
|---|---|
| Message-ID | <CALjSBEvZ_5FU90D8ZaM9QjjPBvDPvgDKr+baTd3qSMyDOUyxhg@mail.gmail.com> |
Hi Lakshimi, I'm glad to hear you've resolved the issue. If you're comfortable sharing, would you mind providing some insights into the root cause? On Wed, Sep 18, 2024 at 1:03=E2=80=AFAM Lakshmi Narasimhan Sundararajan <[email protected]> wrote: > > On Fri, Sep 13, 2024 at 7:49=E2=80=AFPM Tony Asleson <[email protected]= > wrote: > > > > You may want to check out https://lwn.net/Articles/457667/ > > > > On Fri, Sep 13, 2024 at 9:06=E2=80=AFAM Zdenek Kabelac <zdenek.kabelac@= gmail.com> wrote: > > > > > > Dne 13. 09. 24 v 7:55 Lakshmi Narasimhan Sundararajan napsal(a): > > > > Hi Ming, > > > > I am still collecting results, so I will present findings that are > > > > confirmed so far. > > > > There is some good news too. > > > > > > > > On Fri, Sep 13, 2024 at 1:11=E2=80=AFAM Ming Hung Tsai <mtsai@redha= t.com> wrote: > > > >> > > > >> Hi, > > > >> > > > >> On Wed, Sep 11, 2024 at 12:05=E2=80=AFAM Lakshmi Narasimhan Sundar= arajan > > > >> <[email protected]> wrote: > > > > > > > My application that is consuming the thin device pumps IO traffic > > > > directly on the raw block device. > > > > My application also keeps a journal record outside the thin pool an= d > > > > after power recycled, reading the data back > > > > did not guarantee sync consistency. > > > > > > > > I wrote a sample program that is trying to recreate this outside my= application. > > > > here it is: sulakshm/iotest: iotest (github.com) > > > > I am still refining it, as I have not seen the problem with this to= ol yet. > > > > But the logic is similar to how my application consumes thin dev; a= nd > > > > my app can reproduce this very easily. > > > > > > > > As I said before, I am still collecting additional information from > > > > many internal tests. > > > > So far, I can see this problem even in 6.5 kernel. > > > > > > > > The latest distro/linux kernel where this problem is seen. > > > >> 6.5.0-15-generic #15~22.04.1-Ubuntu SMP PREEMPT_DYNAMIC Fri Jan 12= 18:54:30 UTC 2 x86_64 x86_64 x86_64 GNU/Linux > > > > > > > > > > > > There is a workaround to this problem that is looking promising, te= sts > > > > ongoing still. > > > > > > > > The sync point from my application is a sync(fd) of the thin dev. > > > > This has proven insufficient. > > > > In addition, I had to perform "dmsetup suspend pool -> dmsetup resu= me pool". > > > > This guarantees sync point consistency. > > > > > > > > > > Hi > > > > > > Not exactly sure what your app is all exactly doing - however there i= s > > > cut&paste from 'fsync()' manpage: > > > > > > --- > > > Calling fsync() does not necessarily ensure that the entry in the dir= ectory > > > containing the file has also reached disk. For that an explici= t > > > fsync() on a file descriptor for the directory is also needed. > > > --- > > > > > > For this purpose our 'test suite' app basically 'opens' whole device= and > > > fsync and close it - to ensure synchronization point flush. > > > > > > Thus 'suspend & resume' of the whole thin device could be possibly un= necessary > > > - just do a fsync() on blockdevice fd. > > > > > > > > > You can also play fun games with 'fsfreeze' operation. > > > Good day all! > > This issue has been rootcaused successfully to an issue with my applicati= on. > It had been a tough last week given the nature of the issue, and > thanks for everyone who reached out with helpful suggestions. > As part of this process and bug verification, I would gladly submit > that the thin pool implementation did withstand the power cycle tests > wonderfully. > > Best regards and you all have a wonderful day. > > > > > > > > Regards > > > > > > Zdenek > > > > > > > > > > > >