Re: [lustre-devel] RFC: Spill device for Lustre OSD

Jinshan Xiong via lustre-devel <[email protected]> Tue, 4 Nov 2025 15:54:32 -0800
Newsgroups org.lustre.lists.lustre-devel
Message-ID <CAEp8vpi9nRVkBUes8-0RE9YFr7AUPUM+L4AU73UcVU_e6t3ehg@mail.gmail.com>
--===============6683782866803050681==
Content-Type: multipart/alternative; boundary="00000000000046f1090642cd8ee1"

--00000000000046f1090642cd8ee1
Content-Type: text/plain; charset="UTF-8"
Content-Transfer-Encoding: quoted-printable

On Tue, Nov 4, 2025 at 3:48=E2=80=AFPM Andreas Dilger <[email protected]> w=
rote:

>
> Timothy Day <[email protected]> wrote:
> >>>> I haven=E2=80=99t seen any mention of failover yet in this conversat=
ion (may
> have missed it), but if the device is truly local, then in failed over
> configurations the data is inaccessible.  If it=E2=80=99s *not* local, wh=
y not just
> make the device part of the OST or an independent OST?
> >>>
> >>> It won't be local. Actually, this is designed for the cloud.
> >>
> >> I don't understand how 'local' is being used. Cloud or not, all of the
> >> Lustre client, servers, and backend storage service will be co-located
> >> in the same data center. I think Patrick is asking whether the spill
> >> device will be physically connected to OSS server, or be provided over
> >> something like SAN? Either way, presenting this device as an independe=
nt
> >> OST brings back the pain of manually managing data placement from the
> >> client - which this design is trying to avoid.
> >>
> >>> We already have tiered storage based on mirroring; however, that stil=
l
> requires clients to move data and a file system level scanner to decide
> which files move to the cold tier. It's cumbersome to maintain those
> clients.
> >>
> >> Agree, it's not ideal.
>
> Regardless of how the spill device is implemented, there will need to be
> some scanning of the front OSD device to find/manage objects to mirror
> and release.  This could be done directly on the OST with something like
> DDN's lipe_find3 utility, or older scanners like lester, zester, e2scan,
> etc. that scan the local ldiskfs block device directly.
>
> If the overhead of a local Lustre mount on the OSS is problematic, that
> seems like something which could/should be fixed?  The local mounts are
> already "non-recoverable" so that they do not get an entry in last_rcvd
> and their absence does not cause any recovery issues.
>
> The main issue we've seen with local mountpoints is that this can confuse
> HA and prevent Lustre module unloading if they are not taken into account
> during cleanup.
>

You're right. That's actually why we didn't do it in the first place. If an
OSS crashes, it will definitely lead to recovery timeout and client
eviction.


>
> Cheers, Andreas
>
>
>
>
>
>

--00000000000046f1090642cd8ee1
Content-Type: text/html; charset="UTF-8"
Content-Transfer-Encoding: quoted-printable

<div dir=3D"ltr"><div dir=3D"ltr"><br></div><br><div class=3D"gmail_quote g=
mail_quote_container"><div dir=3D"ltr" class=3D"gmail_attr">On Tue, Nov 4, =
2025 at 3:48=E2=80=AFPM Andreas Dilger &lt;<a href=3D"mailto:adilger@dilger=
.ca">[email protected]</a>&gt; wrote:<br></div><blockquote class=3D"gmail_q=
uote" style=3D"margin:0px 0px 0px 0.8ex;border-left:1px solid rgb(204,204,2=
04);padding-left:1ex"><br>
Timothy Day &lt;<a href=3D"mailto:[email protected]" target=3D"_blank">timd=
[email protected]</a>&gt; wrote:<br>
&gt;&gt;&gt;&gt; I haven=E2=80=99t seen any mention of failover yet in this=
 conversation (may have missed it), but if the device is truly local, then =
in failed over configurations the data is inaccessible.=C2=A0 If it=E2=80=
=99s *not* local, why not just make the device part of the OST or an indepe=
ndent OST?<br>
&gt;&gt;&gt; <br>
&gt;&gt;&gt; It won&#39;t be local. Actually, this is designed for the clou=
d.<br>
&gt;&gt; <br>
&gt;&gt; I don&#39;t understand how &#39;local&#39; is being used. Cloud or=
 not, all of the<br>
&gt;&gt; Lustre client, servers, and backend storage service will be co-loc=
ated<br>
&gt;&gt; in the same data center. I think Patrick is asking whether the spi=
ll<br>
&gt;&gt; device will be physically connected to OSS server, or be provided =
over<br>
&gt;&gt; something like SAN? Either way, presenting this device as an indep=
endent<br>
&gt;&gt; OST brings back the pain of manually managing data placement from =
the<br>
&gt;&gt; client - which this design is trying to avoid.<br>
&gt;&gt; <br>
&gt;&gt;&gt; We already have tiered storage based on mirroring; however, th=
at still requires clients to move data and a file system level scanner to d=
ecide which files move to the cold tier. It&#39;s cumbersome to maintain th=
ose clients.<br>
&gt;&gt; <br>
&gt;&gt; Agree, it&#39;s not ideal.<br>
<br>
Regardless of how the spill device is implemented, there will need to be<br=
>
some scanning of the front OSD device to find/manage objects to mirror<br>
and release.=C2=A0 This could be done directly on the OST with something li=
ke<br>
DDN&#39;s lipe_find3 utility, or older scanners like lester, zester, e2scan=
,<br>
etc. that scan the local ldiskfs block device directly.<br>
<br>
If the overhead of a local Lustre mount on the OSS is problematic, that<br>
seems like something which could/should be fixed?=C2=A0 The local mounts ar=
e<br>
already &quot;non-recoverable&quot; so that they do not get an entry in las=
t_rcvd<br>
and their absence does not cause any recovery issues.<br>
<br>
The main issue we&#39;ve seen with local mountpoints is that this can confu=
se<br>
HA and prevent Lustre module unloading if they are not taken into account<b=
r>
during cleanup.<br></blockquote><div><br></div><div>You&#39;re right. That&=
#39;s actually=C2=A0why we didn&#39;t do it in the first place. If an OSS c=
rashes, it will definitely lead to recovery timeout and client eviction.</d=
iv><div>=C2=A0</div><blockquote class=3D"gmail_quote" style=3D"margin:0px 0=
px 0px 0.8ex;border-left:1px solid rgb(204,204,204);padding-left:1ex">
<br>
Cheers, Andreas<br>
<br>
<br>
<br>
<br>
<br>
</blockquote></div></div>

--00000000000046f1090642cd8ee1--

--===============6683782866803050681==
Content-Type: text/plain; charset="us-ascii"
MIME-Version: 1.0
Content-Transfer-Encoding: 7bit
Content-Disposition: inline

_______________________________________________
lustre-devel mailing list
[email protected]
http://lists.lustre.org/listinfo.cgi/lustre-devel-lustre.org

--===============6683782866803050681==--