Re: Seagate ST6000DM003-2CY186 "flapping". Disconnecting and reconnecting randomly.
Daniel Lysfjord <[email protected]>
| Newsgroups | gmane.os.freebsd.questions |
|---|---|
| Message-ID | <[email protected]> |
On 2026-08-11 00:36, Jonathan wrote: > On 2026-08-10 18:25, Daniel Lysfjord wrote: >> On 2026-08-11 00:10, Jonathan wrote: >>>>> >>>>> SMART Attributes Data Structure revision number: 10 >>>>> Vendor Specific SMART Attributes with Thresholds: >>>>> ID# ATTRIBUTE_NAME FLAG VALUE WORST THRESH TYPE >>>>> UPDATED WHEN_FAILED RAW_VALUE >>>>> 1 Raw_Read_Error_Rate 0x000f 073 064 006 Pre-fail >>>>> Always - 20195647 >>>>> 3 Spin_Up_Time 0x0003 098 092 000 Pre-fail >>>>> Always - 0 >>>>> 4 Start_Stop_Count 0x0032 100 100 020 Old_age >>>>> Always - 85 >>>>> 5 Reallocated_Sector_Ct 0x0033 100 100 010 Pre-fail >>>>> Always - 0 >>>>> 7 Seek_Error_Rate 0x000f 070 060 045 Pre-fail >>>>> Always - 10214727 >>>>> 9 Power_On_Hours 0x0032 100 100 000 Old_age >>>>> Always - 726h+57m+29.053s >>>>> 10 Spin_Retry_Count 0x0013 100 100 097 Pre-fail >>>>> Always - 0 >>>>> 12 Power_Cycle_Count 0x0032 100 100 020 Old_age >>>>> Always - 63 >>>>> 183 Runtime_Bad_Block 0x0032 100 100 000 Old_age >>>>> Always - 0 >>>>> 184 End-to-End_Error 0x0032 100 100 099 Old_age >>>>> Always - 0 >>>>> 187 Reported_Uncorrect 0x0032 100 100 000 Old_age >>>>> Always - 0 >>>>> 188 Command_Timeout 0x0032 100 100 000 Old_age >>>>> Always - 0 0 2 >>>>> 189 High_Fly_Writes 0x003a 100 100 000 Old_age >>>>> Always - 0 >>>>> 190 Airflow_Temperature_Cel 0x0022 060 044 040 Old_age >>>>> Always - 40 (Min/Max 39/40) >>>>> 191 G-Sense_Error_Rate 0x0032 100 100 000 Old_age >>>>> Always - 0 >>>>> 192 Power-Off_Retract_Count 0x0032 100 100 000 Old_age >>>>> Always - 94 >>>>> 193 Load_Cycle_Count 0x0032 100 100 000 Old_age >>>>> Always - 315 >>>>> 194 Temperature_Celsius 0x0022 040 056 000 Old_age >>>>> Always - 40 (0 16 0 0 0) >>>>> 195 Hardware_ECC_Recovered 0x001a 073 064 000 Old_age >>>>> Always - 20195647 >>>>> 197 Current_Pending_Sector 0x0012 100 100 000 Old_age >>>>> Always - 0 >>>>> 198 Offline_Uncorrectable 0x0010 100 100 000 Old_age >>>>> Offline - 0 >>>>> 199 UDMA_CRC_Error_Count 0x003e 200 200 000 Old_age >>>>> Always - 0 >>>>> 240 Head_Flying_Hours 0x0000 100 253 000 Old_age >>>>> Offline - 237h+57m+03.832s >>>>> 241 Total_LBAs_Written 0x0000 100 253 000 Old_age >>>>> Offline - 8646277300 >>>>> 242 Total_LBAs_Read 0x0000 100 253 000 Old_age >>>>> Offline - 888924720 >> >> SMART data is not the most reliable to tell anything about anything, >> but, are the Raw_Read_Error_Rate, Seek_Error_Rate and >> Hardware_ECC_Recovered number normal for an SMR drive? It's been ages >> since I last had a spinning rust attached, so, I picked up one of my >> old Seagate Barracudas.. >> >> 1 Raw_Read_Error_Rate 0x000f 101 087 006 Pre-fail >> Always - 3467552 >> 7 Seek_Error_Rate 0x000f 089 060 030 Pre-fail >> Always - 809092981 >> 9 Power_On_Hours 0x0032 044 044 000 Old_age >> Always - 49295 >> 195 Hardware_ECC_Recovered 0x001a 021 003 000 Old_age >> Always - 3467552 >> >> To my eyes, it seems like your drive isn't too happy about something. >> I've seen disks become unhappy due to excess vibration.. Combining >> vibration, SMR and consumer drive might cause the firmware to reset >> the drive instead of timing out? > > I wondered about that as well. None of my other non-SMR drives have > anything close to those numbers. Unfortunately I have no idea how to > tell if that's normal or how to diagnose that as the issue. The drive > is laying next to the server right now so the only vibration it should > be experiencing is its own and it's still randomly dropping off the > controller. At least that what I think is happening based on the > "detached" message in dmesg. It is a drive that was in an external USB > enclosure and I extracted it. I wonder if Seagate did some nasty "reset > randomly if not in the enclosure" thing, I really hope not. > > -- > Jonathan Seagate are stupid, but not Apple-level of stupid? Have you tried plotting all the data you can get about your storage system (netdata might be the simplest thing to set up to see all the things over time), to see if there's a corelation to something? I would've tossed the drive as soon as it starts causing problems. A drive dropping offline (and, hence, triggering higher load on other drives when it needs to be resilvered) isn't really helping your array. It might do better as (one of more than one) target drive for backups, rather than having it in an active array.. Just my €0.2