Re: [lustre-devel] RFC: Spill device for Lustre OSD
Jinshan Xiong via lustre-devel <[email protected]> Tue, 4 Nov 2025 09:51:44 -0800
| Newsgroups | org.lustre.lists.lustre-devel |
|---|---|
| Message-ID | <CAEp8vpjcGobjG-vrHihY3mPB7Lq7a2Hv+-FSPJUVY6R2_w+h_A@mail.gmail.com> |
--===============2101585686254843481== Content-Type: multipart/alternative; boundary="000000000000de642b0642c87c23" --000000000000de642b0642c87c23 Content-Type: text/plain; charset="UTF-8" Content-Transfer-Encoding: quoted-printable On Tue, Nov 4, 2025 at 8:37=E2=80=AFAM Andreas Dilger <[email protected]> wro= te: > On Nov 3, 2025, at 18:58, Oleg Drokin via lustre-devel < > [email protected]> wrote: > > > On Mon, 2025-11-03 at 16:33 -0800, Jinshan Xiong wrote: > > I guess users won't have 1PB OSTs, will they? > > > There probably are already? NASA has a known 0.5P OST configuration: > > https://www.nas.nasa.gov/hecc/support/kb/lustre-progressive-file-layout-(= pfl)-with-ssd-and-hdd-pools_680.html#:~:text=3DThe%20available%20SSD%20spac= e%20in%20each%20filesystem,decimal%20(far%20right)%20labels%20of%20each%20O= ST > > > In order to maximize rebuild performance for declustered parity RAID, > there are OSTs in production with 90x20TB HDDs =3D 1.4 PB today, > and requests to have even larger OSTs. We've done a bunch of work > to improve huge ldiskfs OST performance, including the hybrid OST > patches like https://review.whamcloud.com/51625 ("LU-16750 ldiskfs: > optimize metadata allocation for hybrid LUNs"), but there could still > be further improvements in supporting such large OSTs. > Since they are already HDDs, this feature won't apply because it makes no sense for them. For SSDs, cloud users prefer to spread them across multiple servers to fully utilize disk bandwidth. Bandwidth is likely the first priority. Once that's solved, they want larger capacity so they don't have to load new worksets each time. > > Cheers, Andreas > =E2=80=94 > Andreas Dilger > Lustre Principal Architect > Whamcloud/DDN > > > > > --000000000000de642b0642c87c23 Content-Type: text/html; charset="UTF-8" Content-Transfer-Encoding: quoted-printable <div dir=3D"ltr"><div dir=3D"ltr"><br></div><br><div class=3D"gmail_quote g= mail_quote_container"><div dir=3D"ltr" class=3D"gmail_attr">On Tue, Nov 4, = 2025 at 8:37=E2=80=AFAM Andreas Dilger <<a href=3D"mailto:[email protected]= m">[email protected]</a>> wrote:<br></div><blockquote class=3D"gmail_quote= " style=3D"margin:0px 0px 0px 0.8ex;border-left:1px solid rgb(204,204,204);= padding-left:1ex"> <div> On Nov 3, 2025, at 18:58, Oleg Drokin via lustre-devel <<a href=3D"mailt= o:[email protected]" target=3D"_blank">[email protected]= e.org</a>> wrote:<br> <div> <blockquote type=3D"cite"><br> <div><span style=3D"font-family:Helvetica;font-size:14px;font-style:normal;= font-variant-caps:normal;font-weight:400;letter-spacing:normal;text-align:s= tart;text-indent:0px;text-transform:none;white-space:normal;word-spacing:0p= x;text-decoration:none;float:none;display:inline">On Mon, 2025-11-03 at 16:33 -0800, Jinshan Xiong wrote:</span><br style=3D"fo= nt-family:Helvetica;font-size:14px;font-style:normal;font-variant-caps:norm= al;font-weight:400;letter-spacing:normal;text-align:start;text-indent:0px;t= ext-transform:none;white-space:normal;word-spacing:0px;text-decoration:none= "> <blockquote type=3D"cite" style=3D"font-family:Helvetica;font-size:14px;fon= t-style:normal;font-variant-caps:normal;font-weight:400;letter-spacing:norm= al;text-align:start;text-indent:0px;text-transform:none;white-space:normal;= word-spacing:0px;text-decoration:none"> I guess users won't have 1PB OSTs, will they?<br> </blockquote> <br style=3D"font-family:Helvetica;font-size:14px;font-style:normal;font-va= riant-caps:normal;font-weight:400;letter-spacing:normal;text-align:start;te= xt-indent:0px;text-transform:none;white-space:normal;word-spacing:0px;text-= decoration:none"> <span style=3D"font-family:Helvetica;font-size:14px;font-style:normal;font-= variant-caps:normal;font-weight:400;letter-spacing:normal;text-align:start;= text-indent:0px;text-transform:none;white-space:normal;word-spacing:0px;tex= t-decoration:none;float:none;display:inline">There probably are already? NASA has a known 0.5P OST configuration:</span><br s= tyle=3D"font-family:Helvetica;font-size:14px;font-style:normal;font-variant= -caps:normal;font-weight:400;letter-spacing:normal;text-align:start;text-in= dent:0px;text-transform:none;white-space:normal;word-spacing:0px;text-decor= ation:none"> <a href=3D"https://www.nas.nasa.gov/hecc/support/kb/lustre-progressive-file= -layout-(pfl)-with-ssd-and-hdd-pools_680.html#:~:text=3DThe%20available%20S= SD%20space%20in%20each%20filesystem,decimal%20(far%20right)%20labels%20of%2= 0each%20OST" style=3D"font-family:Helvetica;font-size:14px;font-style:norma= l;font-variant-caps:normal;font-weight:400;letter-spacing:normal;text-align= :start;text-indent:0px;text-transform:none;white-space:normal;word-spacing:= 0px" target=3D"_blank">https://www.nas.nasa.gov/hecc/support/kb/lustre-prog= ressive-file-layout-(pfl)-with-ssd-and-hdd-pools_680.html#:~:text=3DThe%20a= vailable%20SSD%20space%20in%20each%20filesystem,decimal%20(far%20right)%20l= abels%20of%20each%20OST</a><br style=3D"font-family:Helvetica;font-size:14p= x;font-style:normal;font-variant-caps:normal;font-weight:400;letter-spacing= :normal;text-align:start;text-indent:0px;text-transform:none;white-space:no= rmal;word-spacing:0px;text-decoration:none"> </div> </blockquote> </div> <div><br> </div> In order to maximize rebuild performance for declustered parity RAID, <div>there are OSTs in production with 90x20TB HDDs =3D 1.4 PB today,</div> <div>and requests to have even larger OSTs.=C2=A0 We've done a bunch of= work</div> <div>to improve huge ldiskfs OST performance, including the hybrid OST</div= > <div>patches like=C2=A0<a href=3D"https://review.whamcloud.com/51625" targe= t=3D"_blank">https://review.whamcloud.com/51625</a>=C2=A0("LU-16750 ld= iskfs:</div> <div>optimize metadata allocation for hybrid LUNs"), but there could s= till</div> <div>be further improvements in supporting such large OSTs.<br></div></div>= </blockquote><div><br></div><div>Since they are already HDDs, this feature = won't apply because it makes no sense for them.</div><div><br></div><di= v>For SSDs, cloud users prefer to spread them across multiple servers to fu= lly utilize disk bandwidth. Bandwidth is likely the first priority. Once th= at's solved, they want larger capacity so they don't have to load n= ew worksets each time.=C2=A0</div><div><br></div><div>=C2=A0</div><blockquo= te class=3D"gmail_quote" style=3D"margin:0px 0px 0px 0.8ex;border-left:1px = solid rgb(204,204,204);padding-left:1ex"><div><div> <div><br> <div> <div dir=3D"auto" style=3D"color:rgb(0,0,0);letter-spacing:normal;text-alig= n:start;text-indent:0px;text-transform:none;white-space:normal;word-spacing= :0px;text-decoration:none"> <div dir=3D"auto" style=3D"color:rgb(0,0,0);letter-spacing:normal;text-alig= n:start;text-indent:0px;text-transform:none;white-space:normal;word-spacing= :0px;text-decoration:none"> <div>Cheers, Andreas</div> <div>=E2=80=94</div> <div>Andreas Dilger</div> <div>Lustre Principal Architect</div> <div>Whamcloud/DDN</div> </div> <br> </div> <br> <br> </div> <br> </div> </div> </div> </blockquote></div></div> --000000000000de642b0642c87c23-- --===============2101585686254843481== Content-Type: text/plain; charset="us-ascii" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit Content-Disposition: inline _______________________________________________ lustre-devel mailing list [email protected] http://lists.lustre.org/listinfo.cgi/lustre-devel-lustre.org --===============2101585686254843481==--