Array inexplicably degrades itself, but no stale objects to remove
"Ryan Churches" <[email protected]> Thu, 27 Dec 2007 02:21:03 -0500
| Newsgroups | gmane.linux.evms.devel |
|---|---|
| Message-ID | <[email protected]> |
--===============0547338940==
Content-Type: multipart/alternative;
boundary="----=_Part_19716_19759245.1198740063461"
------=_Part_19716_19759245.1198740063461
Content-Type: text/plain; charset=ISO-8859-1
Content-Transfer-Encoding: 7bit
Content-Disposition: inline
Every now and then (every few months) my mirroring raid array degrades
itself. I have no idea how it happens, or why. If I'm lucky I can (IIRC)
"remove stale object" then "add object" and voila. Sometimes this is not
the case.
My setup is as follows: Gentoo Linux 64bit SMP. sda1, 2, and 3 are
mirrored to sdb1, 2, and 3. sda2 is also in a lvm2 storage container. This
is the typical setup recommended by the gentoo wiki as of a few months ago.
In the past the only way Ive been able to repair this has been to
essentially back up my root directory and and rebuild it. I am sick of
doing that, and now that I am positive this is just going to keep happening
it might be good to fix the underlying problem.
First, and perhaps not related, when I start evms I always get this mesasge.
LVM2: Object sdb3 has an LVM2 PV label and header, but the recorded size of
the object (75007232 sectors) does not match
the actual size (75007485 sectors). Please indicate whether or not sdb3 is
an LVM2 PV.
If your container includes an MD RAID region, it's possible that LVM2 has
found the PV label on one of that region's
child objects instead of on the MD region itself. If this is the case, then
object sdb3 is most likely NOT one of the
LVM2 PVs.
Choosing "no" here is the default, and is always safe, since no changes will
be made to your configuration. Choosing
"yes" will modify your configuration, and will cause problems if it's not
the correct choice. The only time you would
really need to choose "yes" here is if you are converting an existing
container from using the LVM2 tools to using EVMS,
and the container is NOT created from an MD RAID region. If you created and
manage your containers only with EVMS, you
should always be able to answer "no".
If you answer "no" and your volumes are correctly discovered and activated,
you may disable this message in the future by
editing the EVMS config file and setting the device_size_prompt option to
"no" in the lvm2 section.
The following responses are available:
*1 = No, it is not a PV.
(Then I get the same message for sda3)
Then I get this message
MDRaid1RegMgr: Region md/md1 is currently in degraded mode. To
bring it back to normal state, add 0 new spare device to replace the faulty
or missing device.
MDRaid1RegMgr: Region md/md2 is currently in degraded mode. To bring it
back to normal state, add 0 new spare device to
replace the faulty or missing device.
As Ive said, I have no idea how or why this happen, it just does.
To confirm, when i cat /proc/mdstat
claudia src # cat /proc/mdstat
Personalities : [linear] [raid0] [raid1] [multipath] [faulty]
md2 : active raid1 dm-3[0]
37503616 blocks [2/1] [U_]
md1 : active raid1 dm-7[0]
1052160 blocks [2/1] [U_]
md0 : active raid1 dm-1[1] dm-0[0]
521984 blocks [2/2] [UU]
unused devices: <none>
So, I'm scared to hose my data so I'm backing up. While I do that, I'm
wondering if anyone knows how I can fix this once and for all.
------=_Part_19716_19759245.1198740063461
Content-Type: text/html; charset=ISO-8859-1
Content-Transfer-Encoding: 7bit
Content-Disposition: inline
Every now and then (every few months) my mirroring raid array degrades itself. I have no idea how it happens, or why. If I'm lucky I can (IIRC) "remove stale object" then "add object" and voila. Sometimes this is not the case.
<br><br>My setup is as follows: Gentoo Linux 64bit SMP. sda1, 2, and 3 are mirrored to sdb1, 2, and 3. sda2 is also in a lvm2 storage container. This is the typical setup recommended by the gentoo wiki as of a few months ago.
<br><br>In the past the only way Ive been able to repair this has been to essentially back up my root directory and and rebuild it. I am sick of doing that, and now that I am positive this is just going to keep happening it might be good to fix the underlying problem.
<br><br>First, and perhaps not related, when I start evms I always get this mesasge.<br><br><br><div style="margin-left: 40px;">LVM2: Object sdb3 has an LVM2 PV label and header, but the recorded size of the object (75007232 sectors) does not match
<br>the actual size (75007485 sectors). Please indicate whether or not sdb3 is an LVM2 PV.<br><br>If your container includes an MD RAID region, it's possible that LVM2 has found the PV label on one of that region's
<br>child objects instead of on the MD region itself. If this is the case, then object sdb3 is most likely NOT one of the<br>LVM2 PVs.<br><br>Choosing "no" here is the default, and is always safe, since no changes will be made to your configuration. Choosing
<br>"yes" will modify your configuration, and will cause problems if it's not the correct choice. The only time you would<br>really need to choose "yes" here is if you are converting an existing container from using the LVM2 tools to using EVMS,
<br>and the container is NOT created from an MD RAID region. If you created and manage your containers only with EVMS, you<br>should always be able to answer "no".<br><br>If you answer "no" and your volumes are correctly discovered and activated, you may disable this message in the future by
<br>editing the EVMS config file and setting the device_size_prompt option to "no" in the lvm2 section.<br>The following responses are available:<br>*1 = No, it is not a PV.<br></div><br><br>(Then I get the same message for sda3)
<br><br>Then I get this message<br><br><div style="margin-left: 40px;">MDRaid1RegMgr: Region md/md1 is currently in degraded mode. To<br>bring it back to normal state, add 0 new spare device to replace the faulty or missing device.
<br><br>MDRaid1RegMgr: Region md/md2 is currently in degraded mode. To bring it back to normal state, add 0 new spare device to<br>replace the faulty or missing device.<br></div><br>As Ive said, I have no idea how or why this happen, it just does.
<br><br>To confirm, when i cat /proc/mdstat<br><br><div style="margin-left: 40px;">claudia src # cat /proc/mdstat<br>Personalities : [linear] [raid0] [raid1] [multipath] [faulty]<br>md2 : active raid1 dm-3[0]<br> 37503616 blocks [2/1] [U_]
<br><br>md1 : active raid1 dm-7[0]<br> 1052160 blocks [2/1] [U_]<br><br>md0 : active raid1 dm-1[1] dm-0[0]<br> 521984 blocks [2/2] [UU]<br><br>unused devices: <none><br><br></div>So, I'm scared to hose my data so I'm backing up. While I do that, I'm wondering if anyone knows how I can fix this once and for all.
<br><br><br>
------=_Part_19716_19759245.1198740063461--
--===============0547338940==
Content-Type: text/plain; charset="us-ascii"
MIME-Version: 1.0
Content-Transfer-Encoding: 7bit
Content-Disposition: inline
-------------------------------------------------------------------------
This SF.net email is sponsored by: Microsoft
Defy all challenges. Microsoft(R) Visual Studio 2005.
http://clk.atdmt.com/MRT/go/vse0120000070mrt/direct/01/
--===============0547338940==
Content-Type: text/plain; charset="us-ascii"
MIME-Version: 1.0
Content-Transfer-Encoding: 7bit
Content-Disposition: inline
_______________________________________________
Evms-devel mailing list
[email protected]
To subscribe/unsubscribe, please visit:
https://lists.sourceforge.net/lists/listinfo/evms-devel
--===============0547338940==--