Two node MetroCluster performance issues?
Heino Walther <[email protected]> Tue, 13 Dec 2022 21:12:58 +0000
| Newsgroups | gmane.comp.hardware.netapp |
|---|---|
| Message-ID | <PR3PR09MB5393F70EFD6138BA359C4D32B8E39@PR3PR09MB5393.eurprd09.prod.outlook.com> |
--===============3655035678152808527==
Content-Language: da-DK
Content-Type: multipart/alternative;
boundary="_000_PR3PR09MB5393F70EFD6138BA359C4D32B8E39PR3PR09MB5393eurp_"
--_000_PR3PR09MB5393F70EFD6138BA359C4D32B8E39PR3PR09MB5393eurp_
Content-Type: text/plain; charset="Windows-1252"
Content-Transfer-Encoding: quoted-printable
Hi there
I have a setup with a two node Fabric MetroCluster with A300 nodes and four=
Brocade 6510 switches places about 10KM from each other.
Our ESXi 7.0 hosts connects via 16G FC using another four frontend Brocade =
6510 switches.
Our ESXi hosts can see four paths for each LUN they are presented from the =
SVMs.
ESXi show all four paths as Active (I/O) which I find a bit odd, because tw=
o of the paths are remote to the ESXi hosts=85
Both the ESXi and the igroup on the NetApp is configured for ALUA=85 but I =
have googled that there can be issues with this, and maybe we have those is=
sues=85
The main issue is that if we do a few performance tests across different da=
tastores (local and remote to the ESXi host), we see OK performance for the=
local datastores (800+MB/sec.) but if we try to test against a datastore t=
hat is remote to the ESXi host we see 50-60MB/sec. which is a huge differen=
ce and leads us to question the setup=85
We are aware that especially writing to a remote datastore will involve a t=
ransfer to the remote controller (10KM) this controller then has to write t=
his to it=92s peer (remote) controller (10KM) the remote peer than has to s=
end an ack. back to the peer (10KM) which then sends an ack. to the ESXi ho=
st (10KM) so all in all 40KM plus waiting for the systems=85 it is not ide=
al, but is this kind of performance normal?
The two nodes does have an ethernet based cluster interconnect which is cur=
rently linked at 1Gb, and I am beginning to suspect that the data between t=
he two nodes is going via this link? But the more I think about it, the mo=
re it does not make sense?
Our FC interswitch links are running 8Gb, but on the switches we see nothin=
g near saturation of any ports=85 and we of cause also checked for port err=
ors of any kind=85
If anyone has a similar setup, any help would be great=85 we are doing a fe=
w more tests, but we are close to opening a case with NetApp=85
/B
--_000_PR3PR09MB5393F70EFD6138BA359C4D32B8E39PR3PR09MB5393eurp_
Content-Type: text/html; charset="Windows-1252"
Content-Transfer-Encoding: quoted-printable
<html xmlns:o=3D"urn:schemas-microsoft-com:office:office" xmlns:w=3D"urn:sc=
hemas-microsoft-com:office:word" xmlns:m=3D"http://schemas.microsoft.com/of=
fice/2004/12/omml" xmlns=3D"http://www.w3.org/TR/REC-html40">
<head>
<meta http-equiv=3D"Content-Type" content=3D"text/html; charset=3DWindows-1=
252">
<meta name=3D"Generator" content=3D"Microsoft Word 15 (filtered medium)">
<style><!--
/* Font Definitions */
@font-face
{font-family:"Cambria Math";
panose-1:2 4 5 3 5 4 6 3 2 4;}
@font-face
{font-family:Calibri;
panose-1:2 15 5 2 2 2 4 3 2 4;}
/* Style Definitions */
p.MsoNormal, li.MsoNormal, div.MsoNormal
{margin:0cm;
font-size:11.0pt;
font-family:"Calibri",sans-serif;
mso-ligatures:standardcontextual;
mso-fareast-language:EN-US;}
span.EmailStyle17
{mso-style-type:personal-compose;
font-family:"Calibri",sans-serif;
color:windowtext;}
.MsoChpDefault
{mso-style-type:export-only;
font-family:"Calibri",sans-serif;
mso-ligatures:standardcontextual;
mso-fareast-language:EN-US;}
@page WordSection1
{size:612.0pt 792.0pt;
margin:3.0cm 2.0cm 3.0cm 2.0cm;}
div.WordSection1
{page:WordSection1;}
--></style>
</head>
<body lang=3D"DA" link=3D"#0563C1" vlink=3D"#954F72" style=3D"word-wrap:bre=
ak-word">
<div class=3D"WordSection1">
<p class=3D"MsoNormal"><span lang=3D"EN-US">Hi there<o:p></o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-US"><o:p> </o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-US">I have a setup with a two node =
Fabric MetroCluster with A300 nodes and four Brocade 6510 switches places a=
bout 10KM from each other.<o:p></o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-US">Our ESXi 7.0 hosts connects via=
16G FC using another four frontend Brocade 6510 switches.<o:p></o:p></span=
></p>
<p class=3D"MsoNormal"><span lang=3D"EN-US">Our ESXi hosts can see four pat=
hs for each LUN they are presented from the SVMs.<o:p></o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-US">ESXi show all four paths as Act=
ive (I/O) which I find a bit odd, because two of the paths are remote to th=
e ESXi hosts=85<o:p></o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-US">Both the ESXi and the igroup on=
the NetApp is configured for ALUA=85 but I have googled that there can be =
issues with this, and maybe we have those issues=85<o:p></o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-US"><o:p> </o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-US">The main issue is that if we do=
a few performance tests across different datastores (local and remote to t=
he ESXi host), we see OK performance for the local datastores (800+MB/sec.)=
but if we try to test against a datastore
that is remote to the ESXi host we see 50-60MB/sec. which is a huge differ=
ence and leads us to question the setup=85<o:p></o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-US"><o:p> </o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-US">We are aware that especially wr=
iting to a remote datastore will involve a transfer to the remote controlle=
r (10KM) this controller then has to write this to it=92s peer (remote) con=
troller (10KM) the remote peer than has
to send an ack. back to the peer (10KM) which then sends an ack. to the ES=
Xi host (10KM) so all in all 40KM plus waiting for the systems=85 it =
is not ideal, but is this kind of performance normal?<o:p></o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-US">The two nodes does have an ethe=
rnet based cluster interconnect which is currently linked at 1Gb, and I am =
beginning to suspect that the data between the two nodes is going via this =
link? But the more I think about it,
the more it does not make sense?<o:p></o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-US"><o:p> </o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-US">Our FC interswitch links are ru=
nning 8Gb, but on the switches we see nothing near saturation of any ports=
=85 and we of cause also checked for port errors of any kind=85
<o:p></o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-US"><o:p> </o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-US">If anyone has a similar setup, =
any help would be great=85 we are doing a few more tests, but we are close =
to opening a case with NetApp=85<o:p></o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-US"><o:p> </o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-US">/B<o:p></o:p></span></p>
</div>
</body>
</html>
--_000_PR3PR09MB5393F70EFD6138BA359C4D32B8E39PR3PR09MB5393eurp_--
--===============3655035678152808527==
Content-Type: text/plain; charset="us-ascii"
MIME-Version: 1.0
Content-Transfer-Encoding: 7bit
Content-Disposition: inline
_______________________________________________
Toasters mailing list
[email protected]
https://www.teaparty.net/mailman/listinfo/toasters
--===============3655035678152808527==--