Re: [lustre-devel] RFC: Spill device for Lustre OSD
Jinshan Xiong via lustre-devel <[email protected]> Tue, 4 Nov 2025 15:54:32 -0800
| Newsgroups | org.lustre.lists.lustre-devel |
|---|---|
| Message-ID | <CAEp8vpi9nRVkBUes8-0RE9YFr7AUPUM+L4AU73UcVU_e6t3ehg@mail.gmail.com> |
--===============6683782866803050681== Content-Type: multipart/alternative; boundary="00000000000046f1090642cd8ee1" --00000000000046f1090642cd8ee1 Content-Type: text/plain; charset="UTF-8" Content-Transfer-Encoding: quoted-printable On Tue, Nov 4, 2025 at 3:48=E2=80=AFPM Andreas Dilger <[email protected]> w= rote: > > Timothy Day <[email protected]> wrote: > >>>> I haven=E2=80=99t seen any mention of failover yet in this conversat= ion (may > have missed it), but if the device is truly local, then in failed over > configurations the data is inaccessible. If it=E2=80=99s *not* local, wh= y not just > make the device part of the OST or an independent OST? > >>> > >>> It won't be local. Actually, this is designed for the cloud. > >> > >> I don't understand how 'local' is being used. Cloud or not, all of the > >> Lustre client, servers, and backend storage service will be co-located > >> in the same data center. I think Patrick is asking whether the spill > >> device will be physically connected to OSS server, or be provided over > >> something like SAN? Either way, presenting this device as an independe= nt > >> OST brings back the pain of manually managing data placement from the > >> client - which this design is trying to avoid. > >> > >>> We already have tiered storage based on mirroring; however, that stil= l > requires clients to move data and a file system level scanner to decide > which files move to the cold tier. It's cumbersome to maintain those > clients. > >> > >> Agree, it's not ideal. > > Regardless of how the spill device is implemented, there will need to be > some scanning of the front OSD device to find/manage objects to mirror > and release. This could be done directly on the OST with something like > DDN's lipe_find3 utility, or older scanners like lester, zester, e2scan, > etc. that scan the local ldiskfs block device directly. > > If the overhead of a local Lustre mount on the OSS is problematic, that > seems like something which could/should be fixed? The local mounts are > already "non-recoverable" so that they do not get an entry in last_rcvd > and their absence does not cause any recovery issues. > > The main issue we've seen with local mountpoints is that this can confuse > HA and prevent Lustre module unloading if they are not taken into account > during cleanup. > You're right. That's actually why we didn't do it in the first place. If an OSS crashes, it will definitely lead to recovery timeout and client eviction. > > Cheers, Andreas > > > > > > --00000000000046f1090642cd8ee1 Content-Type: text/html; charset="UTF-8" Content-Transfer-Encoding: quoted-printable <div dir=3D"ltr"><div dir=3D"ltr"><br></div><br><div class=3D"gmail_quote g= mail_quote_container"><div dir=3D"ltr" class=3D"gmail_attr">On Tue, Nov 4, = 2025 at 3:48=E2=80=AFPM Andreas Dilger <<a href=3D"mailto:adilger@dilger= .ca">[email protected]</a>> wrote:<br></div><blockquote class=3D"gmail_q= uote" style=3D"margin:0px 0px 0px 0.8ex;border-left:1px solid rgb(204,204,2= 04);padding-left:1ex"><br> Timothy Day <<a href=3D"mailto:[email protected]" target=3D"_blank">timd= [email protected]</a>> wrote:<br> >>>> I haven=E2=80=99t seen any mention of failover yet in this= conversation (may have missed it), but if the device is truly local, then = in failed over configurations the data is inaccessible.=C2=A0 If it=E2=80= =99s *not* local, why not just make the device part of the OST or an indepe= ndent OST?<br> >>> <br> >>> It won't be local. Actually, this is designed for the clou= d.<br> >> <br> >> I don't understand how 'local' is being used. Cloud or= not, all of the<br> >> Lustre client, servers, and backend storage service will be co-loc= ated<br> >> in the same data center. I think Patrick is asking whether the spi= ll<br> >> device will be physically connected to OSS server, or be provided = over<br> >> something like SAN? Either way, presenting this device as an indep= endent<br> >> OST brings back the pain of manually managing data placement from = the<br> >> client - which this design is trying to avoid.<br> >> <br> >>> We already have tiered storage based on mirroring; however, th= at still requires clients to move data and a file system level scanner to d= ecide which files move to the cold tier. It's cumbersome to maintain th= ose clients.<br> >> <br> >> Agree, it's not ideal.<br> <br> Regardless of how the spill device is implemented, there will need to be<br= > some scanning of the front OSD device to find/manage objects to mirror<br> and release.=C2=A0 This could be done directly on the OST with something li= ke<br> DDN's lipe_find3 utility, or older scanners like lester, zester, e2scan= ,<br> etc. that scan the local ldiskfs block device directly.<br> <br> If the overhead of a local Lustre mount on the OSS is problematic, that<br> seems like something which could/should be fixed?=C2=A0 The local mounts ar= e<br> already "non-recoverable" so that they do not get an entry in las= t_rcvd<br> and their absence does not cause any recovery issues.<br> <br> The main issue we've seen with local mountpoints is that this can confu= se<br> HA and prevent Lustre module unloading if they are not taken into account<b= r> during cleanup.<br></blockquote><div><br></div><div>You're right. That&= #39;s actually=C2=A0why we didn't do it in the first place. If an OSS c= rashes, it will definitely lead to recovery timeout and client eviction.</d= iv><div>=C2=A0</div><blockquote class=3D"gmail_quote" style=3D"margin:0px 0= px 0px 0.8ex;border-left:1px solid rgb(204,204,204);padding-left:1ex"> <br> Cheers, Andreas<br> <br> <br> <br> <br> <br> </blockquote></div></div> --00000000000046f1090642cd8ee1-- --===============6683782866803050681== Content-Type: text/plain; charset="us-ascii" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit Content-Disposition: inline _______________________________________________ lustre-devel mailing list [email protected] http://lists.lustre.org/listinfo.cgi/lustre-devel-lustre.org --===============6683782866803050681==--