Re: smartctl output for Crucial 512Gb SSD
Mike <[email protected]>
| Newsgroups | gmane.linux.utilities.smartmontools |
|---|---|
| Message-ID | <CABfcST8_XxPawMgzZk1AgKORDNQgONfr0y+aCEnQJYSn1XA8yg@mail.gmail.com> |
Greetings
Following on from previous message, where running smartctl with the
disk connected via SATA returns a 600PB disk (it is actually 512Gb
Crucial M4), I followed advice from the Crucial message boards and
left the disk idle over night in order to trigger automatic garbade
collection. No change the next morning, below is some smartctl
output. I've tried ddrescue to no avail, reads just fail completely.
user@w530:~/Downloads/smartmontools-6.3$ sudo ./smartctl -x /dev/sdb
[sudo] password for user:
smartctl 6.3 2014-07-26 r3976 [x86_64-linux-3.13.0-34-generic] (local build)
Copyright (C) 2002-14, Bruce Allen, Christian Franke, www.smartmontools.org
=== START OF INFORMATION SECTION ===
Model Family: Crucial/Micron RealSSD m4/C400/P400
Device Model: M4-CT512M4SSD1
Serial Number: 000000001249091FCC21
LU WWN Device Id: 5 00a075 1091fcc21
Firmware Version: 070H
User Capacity: 512,110,190,592 bytes [512 GB]
Sector Size: 512 bytes logical/physical
Rotation Rate: Solid State Device
Form Factor: 2.5 inches
Device is: In smartctl database [for details use: -P show]
ATA Version is: ACS-2, ATA8-ACS T13/1699-D revision 6
SATA Version is: SATA 3.0, 6.0 Gb/s (current: 1.5 Gb/s)
Local Time is: Sat Aug 16 09:52:11 2014 BST
SMART support is: Available - device has SMART capability.
SMART support is: Enabled
AAM feature is: Unavailable
APM level is: 254 (maximum performance)
Rd look-ahead is: Enabled
Write cache is: Enabled
ATA Security is: Disabled, NOT FROZEN [SEC1]
Write SCT (Get) Feature Control Command failed: scsi error medium or
hardware error (serious)
Wt Cache Reorder: Unknown (SCT Feature Control command failed)
=== START OF READ SMART DATA SECTION ===
SMART Status command failed: scsi error medium or hardware error (serious)
SMART overall-health self-assessment test result: PASSED
Warning: This result is based on an Attribute check.
General SMART Values:
Offline data collection status: (0x80) Offline data collection activity
was never started.
Auto Offline Data Collection: Enabled.
Self-test execution status: ( 0) The previous self-test routine completed
without error or no self-test has ever
been run.
Total time to complete Offline
data collection: ( 2380) seconds.
Offline data collection
capabilities: (0x7b) SMART execute Offline immediate.
Auto Offline data collection on/off support.
Suspend Offline collection upon new
command.
Offline surface scan supported.
Self-test supported.
Conveyance Self-test supported.
Selective Self-test supported.
SMART capabilities: (0x0003) Saves SMART data before entering
power-saving mode.
Supports SMART auto save timer.
Error logging capability: (0x01) Error logging supported.
General Purpose Logging supported.
Short self-test routine
recommended polling time: ( 2) minutes.
Extended self-test routine
recommended polling time: ( 39) minutes.
Conveyance self-test routine
recommended polling time: ( 3) minutes.
SCT capabilities: (0x003d) SCT Status supported.
SCT Error Recovery Control supported.
SCT Feature Control supported.
SCT Data Table supported.
SMART Attributes Data Structure revision number: 16
Vendor Specific SMART Attributes with Thresholds:
ID# ATTRIBUTE_NAME FLAGS VALUE WORST THRESH FAIL RAW_VALUE
1 Raw_Read_Error_Rate POSR-K 100 100 050 - 0
5 Reallocated_Sector_Ct PO--CK 100 100 010 - 0
9 Power_On_Hours -O--CK 100 100 001 - 4439
12 Power_Cycle_Count -O--CK 100 100 001 - 515
170 Grown_Failing_Block_Ct PO--CK 100 100 010 - 0
171 Program_Fail_Count -O--CK 100 100 001 - 0
172 Erase_Fail_Count -O--CK 100 100 001 - 0
173 Wear_Leveling_Count PO--CK 100 100 010 - 0
174 Unexpect_Power_Loss_Ct -O--CK 100 100 001 - 398
181 Non4k_Aligned_Access -O---K 100 100 001 - 144 45 99
183 SATA_Iface_Downshift -O--CK 100 100 001 - 0
184 End-to-End_Error PO--CK 100 100 050 - 0
187 Reported_Uncorrect -O--CK 100 100 001 - 0
188 Command_Timeout -O--CK 100 100 001 - 0
189 Factory_Bad_Block_Ct -OSR-- 100 100 001 - 215
194 Temperature_Celsius -O---K 100 100 000 - 0
195 Hardware_ECC_Recovered -O-RCK 100 100 001 - 0
196 Reallocated_Event_Count -O--CK 100 100 001 - 0
197 Current_Pending_Sector -O--CK 100 100 001 - 0
198 Offline_Uncorrectable ----CK 100 100 001 - 0
199 UDMA_CRC_Error_Count -O--CK 100 100 001 - 0
202 Perc_Rated_Life_Used ---RC- 100 100 001 - 0
206 Write_Error_Rate -OSR-- 100 100 001 - 0
||||||_ K auto-keep
|||||__ C event count
||||___ R error rate
|||____ S speed/performance
||_____ O updated online
|______ P prefailure warning
Read SMART Log Directory failed: scsi error medium or hardware error (serious)
General Purpose Log Directory Version 1
Address Access R/W Size Description
0x00 GPL R/O 1 Log Directory
0x03 GPL R/O 16383 Ext. Comprehensive SMART error log
0x04 GPL R/O 255 Device Statistics log
0x07 GPL R/O 3449 Extended self-test log
0x10 GPL R/O 1 NCQ Command Error log
0x11 GPL R/O 1 SATA Phy Event Counters
0x80-0x9f GPL R/W 16 Host vendor specific log
0xa0 GPL VS 2000 Device vendor specific log
0xa1-0xbf GPL VS 1 Device vendor specific log
0xc0 GPL VS 80 Device vendor specific log
0xc1-0xdf GPL VS 1 Device vendor specific log
0xe0 GPL R/W 1 SCT Command/Status
0xe1 GPL R/W 1 SCT Data Transfer
SMART Extended Comprehensive Error Log size 16383 not supported
Read SMART Error Log failed: scsi error medium or hardware error (serious)
SMART Extended Self-test Log size 3449 not supported
Read SMART Self-test Log failed: scsi error medium or hardware error (serious)
Read SMART Selective Self-test Log failed: scsi error medium or
hardware error (serious)
SCT Status Version: 3
SCT Version (vendor specific): 1 (0x0001)
SCT Support Level: 0
Device State: Stand-by (1)
Current Temperature: 0 Celsius
Power Cycle Max Temperature: 0 Celsius
Lifetime Max Temperature: 0 Celsius
SCT Temperature History Version: 2
Temperature Sampling Period: 10 minutes
Temperature Logging Interval: 10 minutes
Min/Max recommended Temperature: 0/70 Celsius
Min/Max Temperature Limit: -5/75 Celsius
Temperature History Size (Index): 478 (38)
Index Estimated Time Temperature Celsius
39 2014-08-13 02:20 ? -
... ..(475 skipped). .. -
37 2014-08-16 09:40 ? -
38 2014-08-16 09:50 0 -
Write SCT (Get) Error Recovery Control Command failed: scsi error
medium or hardware error (serious)
SCT (Get) Error Recovery Control command failed
Device Statistics (GP Log 0x04)
Page Offset Size Value Description
1 ===== = = == General Statistics (rev 2) ==
1 0x008 4 515 Lifetime Power-On Resets
1 0x010 4 4439 Power-on Hours
1 0x018 6 933634265 Logical Sectors Written
1 0x020 6 2708986 Number of Write Commands
1 0x028 6 852431906 Logical Sectors Read
1 0x030 6 6049196 Number of Read Commands
4 ===== = = == General Errors Statistics (rev 1) ==
4 0x008 4 0 Number of Reported Uncorrectable Errors
4 0x010 4 0 Resets Between Cmd Acceptance and Completion
5 ===== = = == Temperature Statistics (rev 1) ==
5 0x008 1 0 Current Temperature
5 0x010 1 0 Average Short Term Temperature
5 0x018 1 0 Average Long Term Temperature
5 0x020 1 0 Highest Temperature
5 0x028 1 0 Lowest Temperature
5 0x030 1 0 Highest Average Short Term Temperature
5 0x038 1 0 Lowest Average Short Term Temperature
5 0x040 1 0 Highest Average Long Term Temperature
5 0x048 1 0 Lowest Average Long Term Temperature
5 0x050 4 - Time in Over-Temperature
5 0x058 1 70 Specified Maximum Operating Temperature
5 0x060 4 - Time in Under-Temperature
5 0x068 1 0 Specified Minimum Operating Temperature
6 ===== = = == Transport Statistics (rev 1) ==
6 0x008 4 0 Number of Hardware Resets
6 0x010 4 0 Number of ASR Events
6 0x018 4 0 Number of Interface CRC Errors
7 ===== = = == Solid State Device Statistics (rev 1) ==
7 0x008 1 21~ Percentage Used Endurance Indicator
|_ ~ normalized value
SATA Phy Event Counters (GP Log 0x11)
ID Size Value Description
0x0001 4 0 Command failed due to ICRC error
0x000a 4 0 Device-to-host register FISes sent due to a COMRESET
On 14 August 2014 20:08, Mike <[email protected]> wrote:
> Hello again
>
> I've put the drive back into the system it was used in, and used the
> one other SATA port available. Running smartctl returns some strange
> output, I've also included the kernel message in case it helps.
>
> root@voyage:~/smartmontools-6.3# ./smartctl -T permissive -x /dev/sda
> smartctl 6.3 2014-07-26 r3976 [i686-linux-3.4.4-voyage] (local build)
> Copyright (C) 2002-14, Bruce Allen, Christian Franke, www.smartmontools.org
>
> === START OF INFORMATION SECTION ===
> Vendor: /3:0:0:0
> Product:
> Compliance: SPC-5
> User Capacity: 600,332,565,813,390,450 bytes [600 PB]
> Logical block size: 774843950 bytes
> scsiModePageOffset: response length too short, resp_len=47 offset=50 bd_len=46
> scsiModePageOffset: response length too short, resp_len=47 offset=50 bd_len=46
>>> Terminate command early due to bad response to IEC mode page
> Read Cache is: Unavailable
> Writeback Cache is: Unavailable
>
> === START OF READ SMART DATA SECTION ===
> Log Sense failed, IE page [scsi response fails sanity test]
> Error Counter logging not supported
>
> scsiModePageOffset: response length too short, resp_len=47 offset=50 bd_len=46
> Device does not support Self Test logging
> Device does not support Background scan results logging
>
>
> [ 3.899820] ata1: SATA max UDMA/133 abar m1024@0xffe40000 port
> 0xffe40100 irq 45
> [ 3.899954] ata2: DUMMY
> [ 3.900067] ata3: SATA max UDMA/133 abar m1024@0xffe40000 port
> 0xffe40200 irq 45
> [ 3.900193] ata4: DUMMY
> [ 4.205145] ata3: SATA link up 1.5 Gbps (SStatus 113 SControl 300)
> [ 4.205279] ata1: SATA link down (SStatus 0 SControl 300)
> [ 4.205686] ata3.00: failed to enable AA (error_mask=0x1)
> [ 4.205772] ata3.00: ATA-9: M4-CT512M4SSD1, 070H, max UDMA/100
> [ 4.205853] ata3.00: 1000215216 sectors, multi 16: LBA48 NCQ (depth 31/32)
> [ 4.206476] ata3.00: failed to enable AA (error_mask=0x1)
> [ 4.206568] ata3.00: configured for UDMA/100 (device error ignored)
> [ 4.207143] scsi 3:0:0:0: Direct-Access ATA M4-CT512M4SSD1
> 070H PQ: 0 ANSI: 5
> [ 4.207977] sd 3:0:0:0: [sda] 1000215216 512-byte logical blocks:
> (512 GB/476 GiB)
> [ 4.208722] sd 3:0:0:0: [sda] Write Protect is off
> [ 4.208822] sd 3:0:0:0: [sda] Mode Sense: 00 3a 00 00
> [ 4.208998] sd 3:0:0:0: [sda] Write cache: enabled, read cache:
> enabled, doesn't support DPO or FUA
> [ 4.569257] scsi 0:0:0:0: Direct-Access TS TS2GUFM-H
> 1100 PQ: 0 ANSI: 0 CCS
> [ 4.571000] sd 0:0:0:0: [sdb] 4014080 512-byte logical blocks:
> (2.05 GB/1.91 GiB)
> [ 4.572153] sd 0:0:0:0: [sdb] Write Protect is off
> [ 4.572260] sd 0:0:0:0: [sdb] Mode Sense: 43 00 00 00
> [ 4.573219] sd 0:0:0:0: [sdb] No Caching mode page present
> [ 4.573323] sd 0:0:0:0: [sdb] Assuming drive cache: write through
> [ 4.577728] sd 0:0:0:0: [sdb] No Caching mode page present
> [ 4.577837] sd 0:0:0:0: [sdb] Assuming drive cache: write through
> [ 4.579449] sdb: sdb1
> [ 4.583344] sd 0:0:0:0: [sdb] No Caching mode page present
> [ 4.583454] sd 0:0:0:0: [sdb] Assuming drive cache: write through
> [ 4.583561] sd 0:0:0:0: [sdb] Attached SCSI disk
> [ 4.587875] sd 3:0:0:0: Attached scsi generic sg0 type 0
> [ 4.588303] sd 0:0:0:0: Attached scsi generic sg1 type 0
> [ 34.720249] ata3.00: exception Emask 0x0 SAct 0x1 SErr 0x0 action 0x6 frozen
> [ 34.720368] ata3.00: failed command: READ FPDMA QUEUED
> [ 34.720482] ata3.00: cmd 60/08:00:00:00:00/00:00:00:00:00/40 tag 0
> ncq 4096 in
> [ 34.720489] res 40/00:00:00:00:00/00:00:00:00:00/00 Emask
> 0x4 (timeout)
> [ 34.720684] ata3.00: status: { DRDY }
> [ 34.720769] ata3: hard resetting link
> [ 40.074131] ata3: link is slow to respond, please be patient (ready=0)
> [ 44.766167] ata3: COMRESET failed (errno=-16)
> [ 44.766267] ata3: hard resetting link
> [ 50.120136] ata3: link is slow to respond, please be patient (ready=0)
> [ 54.812132] ata3: COMRESET failed (errno=-16)
> [ 54.812216] ata3: hard resetting link
> [ 60.166135] ata3: link is slow to respond, please be patient (ready=0)
> [ 89.848141] ata3: COMRESET failed (errno=-16)
> [ 89.848226] ata3: limiting SATA link speed to 1.5 Gbps
> [ 89.848305] ata3: hard resetting link
> [ 94.896060] ata3: COMRESET failed (errno=-16)
> [ 94.896149] ata3: reset failed, giving up
> [ 94.896225] ata3.00: disabled
> [ 94.896300] ata3.00: device reported invalid CHS sector 0
> [ 94.896402] ata3: EH complete
> [ 94.896535] sd 3:0:0:0: [sda] Unhandled error code
> [ 94.896613] sd 3:0:0:0: [sda] Result: hostbyte=0x04 driverbyte=0x00
> [ 94.896726] sd 3:0:0:0: [sda] CDB: cdb[0]=0x28: 28 00 00 00 00 00 00 00 08 00
> [ 94.897243] end_request: I/O error, dev sda, sector 0
> [ 94.897323] Buffer I/O error on device sda, logical block 0
> [ 94.897517] sd 3:0:0:0: [sda] Unhandled error code
> [ 94.897599] sd 3:0:0:0: [sda] Result: hostbyte=0x04 driverbyte=0x00
> [ 94.897712] sd 3:0:0:0: [sda] CDB: cdb[0]=0x28: 28 00 00 00 00 00 00 00 08 00
> [ 94.898216] end_request: I/O error, dev sda, sector 0
> [ 94.898293] Buffer I/O error on device sda, logical block 0
> [ 94.898490] sd 3:0:0:0: [sda] Unhandled error code
> [ 94.898568] sd 3:0:0:0: [sda] Result: hostbyte=0x04 driverbyte=0x00
> [ 94.898681] sd 3:0:0:0: [sda] CDB: cdb[0]=0x28: 28 00 00 00 00 00 00 00 08 00
> [ 94.899194] end_request: I/O error, dev sda, sector 0
> [ 94.899271] Buffer I/O error on device sda, logical block 0
> [ 94.899383] ldm_validate_partition_table(): Disk read failed.
> [ 94.899542] sd 3:0:0:0: [sda] Unhandled error code
> [ 94.899620] sd 3:0:0:0: [sda] Result: hostbyte=0x04 driverbyte=0x00
> [ 94.899732] sd 3:0:0:0: [sda] CDB: cdb[0]=0x28: 28 00 00 00 00 00 00 00 08 00
> [ 94.900237] end_request: I/O error, dev sda, sector 0
> [ 94.900313] Buffer I/O error on device sda, logical block 0
> [ 94.900506] sd 3:0:0:0: [sda] Unhandled error code
> [ 94.900584] sd 3:0:0:0: [sda] Result: hostbyte=0x04 driverbyte=0x00
> [ 94.900697] sd 3:0:0:0: [sda] CDB: cdb[0]=0x28: 28 00 00 00 00 00 00 00 08 00
> [ 94.901210] end_request: I/O error, dev sda, sector 0
> [ 94.901287] Buffer I/O error on device sda, logical block 0
> [ 94.901399] sda: unable to read partition table
> [ 94.902193] sd 3:0:0:0: [sda] READ CAPACITY(16) failed
> [ 94.902281] sd 3:0:0:0: [sda] Result: hostbyte=0x04 driverbyte=0x00
> [ 94.902398] sd 3:0:0:0: [sda] Sense not available.
> [ 94.902682] sd 3:0:0:0: [sda] READ CAPACITY failed
> [ 94.902762] sd 3:0:0:0: [sda] Result: hostbyte=0x04 driverbyte=0x00
> [ 94.902877] sd 3:0:0:0: [sda] Sense not available.
> [ 94.903338] sd 3:0:0:0: [sda] Got wrong page
> [ 94.903419] sd 3:0:0:0: [sda] Assuming drive cache: write through
> [ 94.903501] sd 3:0:0:0: [sda] Attached SCSI disk
>
> On 13 August 2014 22:27, Volker Kuhlmann <[email protected]> wrote:
>> On Wed 13 Aug 2014 03:36:23 NZST +1200, Mike wrote:
>>
>>> My m4 512Gb CT512M4SSD1 suddenly failed to read. Below is the smartctl
>>> output that returns eventually, taking at least a couple of minutes. I've
>>> tried 2 different USB caddys and also SATA connection to laptop running
>>> livecd. The drive has exclusively been mounted read-only on a Linux system
>>> with an ext4 filesystem, and has had unexpected power loss routinely
>>> throughout its life. The values in the attribute table look OK to me I
>>> think, but the serious errors at the bottom look bad. Any ideas greatly
>>> appreciated:
>>
>>> Read SMART Log Directory failed.
>>>
>>> Error SMART Error Log Read failed: scsi error medium or hardware error
>>> (serious)
>>> Smartctl: SMART Error Log Read Failed
>>> Error SMART Error Self-Test Log Read failed: scsi error medium or hardware
>>> error (serious)
>>> Smartctl: SMART Self Test Log Read Failed
>>> Error SMART Read Selective Self-Test Log failed: scsi error medium or
>>> hardware error (serious)
>>> Smartctl: SMART Selective Self Test Log Read Failed
>>
>> This looks like smartctl may not be able to communicate with the drive
>> properly, which is what you expect if the drive/computer interface
>> circuitry is damaged. However, these errors don't seem to exclude the
>> possibility of a serious platter failure - or rather the equivalent in
>> solid state.
>>
>> Moving the disk around between enclosures etc can easily damage the
>> drive electronics if sufficient ESD protection procedures have not been
>> followed (those procedures must be sufficient regardless of your opinion
>> of them ;-) ). If the problem first occurred without having moved any
>> hardware around then the problem is with the drive. If you can reliably
>> exclude the problem being located at the computer, enclosure or
>> connections/cables then it's also the drive.
>>
>> If the problem is with the drive then the drive is most certainly
>> finished. Copy off all the data you can asap and use your backups for
>> the rest.
>>
>> It might be a good idea to treat SSDs the same as spinning versions. The
>> differences are in speed and behaviour after being dropped, not
>> necessarily in long-term reliability of data.
>>
>> HTH,
>>
>> Volker
>>
>> --
>> Volker Kuhlmann
>> http://volker.top.geek.nz/ Please do not CC list postings to me.
>>
>> ------------------------------------------------------------------------------
>> _______________________________________________
>> Smartmontools-support mailing list
>> [email protected]
>> https://lists.sourceforge.net/lists/listinfo/smartmontools-support
------------------------------------------------------------------------------