Re: smartctl output for Crucial 512Gb SSD

Mike <[email protected]>
Newsgroups gmane.linux.utilities.smartmontools
Message-ID <CABfcST8_XxPawMgzZk1AgKORDNQgONfr0y+aCEnQJYSn1XA8yg@mail.gmail.com>
Greetings

Following on from previous message, where running smartctl with the
disk connected via SATA returns a 600PB disk (it is actually 512Gb
Crucial M4), I followed advice from the Crucial message boards and
left the disk idle over night in order to trigger automatic garbade
collection.  No change the next morning, below is some smartctl
output.  I've tried ddrescue to no avail, reads just fail completely.

user@w530:~/Downloads/smartmontools-6.3$ sudo ./smartctl -x /dev/sdb
[sudo] password for user:
smartctl 6.3 2014-07-26 r3976 [x86_64-linux-3.13.0-34-generic] (local build)
Copyright (C) 2002-14, Bruce Allen, Christian Franke, www.smartmontools.org

=== START OF INFORMATION SECTION ===
Model Family:     Crucial/Micron RealSSD m4/C400/P400
Device Model:     M4-CT512M4SSD1
Serial Number:    000000001249091FCC21
LU WWN Device Id: 5 00a075 1091fcc21
Firmware Version: 070H
User Capacity:    512,110,190,592 bytes [512 GB]
Sector Size:      512 bytes logical/physical
Rotation Rate:    Solid State Device
Form Factor:      2.5 inches
Device is:        In smartctl database [for details use: -P show]
ATA Version is:   ACS-2, ATA8-ACS T13/1699-D revision 6
SATA Version is:  SATA 3.0, 6.0 Gb/s (current: 1.5 Gb/s)
Local Time is:    Sat Aug 16 09:52:11 2014 BST
SMART support is: Available - device has SMART capability.
SMART support is: Enabled
AAM feature is:   Unavailable
APM level is:     254 (maximum performance)
Rd look-ahead is: Enabled
Write cache is:   Enabled
ATA Security is:  Disabled, NOT FROZEN [SEC1]
Write SCT (Get) Feature Control Command failed: scsi error medium or
hardware error (serious)
Wt Cache Reorder: Unknown (SCT Feature Control command failed)

=== START OF READ SMART DATA SECTION ===
SMART Status command failed: scsi error medium or hardware error (serious)
SMART overall-health self-assessment test result: PASSED
Warning: This result is based on an Attribute check.

General SMART Values:
Offline data collection status:  (0x80) Offline data collection activity
was never started.
Auto Offline Data Collection: Enabled.
Self-test execution status:      (   0) The previous self-test routine completed
without error or no self-test has ever
been run.
Total time to complete Offline
data collection: ( 2380) seconds.
Offline data collection
capabilities: (0x7b) SMART execute Offline immediate.
Auto Offline data collection on/off support.
Suspend Offline collection upon new
command.
Offline surface scan supported.
Self-test supported.
Conveyance Self-test supported.
Selective Self-test supported.
SMART capabilities:            (0x0003) Saves SMART data before entering
power-saving mode.
Supports SMART auto save timer.
Error logging capability:        (0x01) Error logging supported.
General Purpose Logging supported.
Short self-test routine
recommended polling time: (   2) minutes.
Extended self-test routine
recommended polling time: (  39) minutes.
Conveyance self-test routine
recommended polling time: (   3) minutes.
SCT capabilities:       (0x003d) SCT Status supported.
SCT Error Recovery Control supported.
SCT Feature Control supported.
SCT Data Table supported.

SMART Attributes Data Structure revision number: 16
Vendor Specific SMART Attributes with Thresholds:
ID# ATTRIBUTE_NAME          FLAGS    VALUE WORST THRESH FAIL RAW_VALUE
  1 Raw_Read_Error_Rate     POSR-K   100   100   050    -    0
  5 Reallocated_Sector_Ct   PO--CK   100   100   010    -    0
  9 Power_On_Hours          -O--CK   100   100   001    -    4439
 12 Power_Cycle_Count       -O--CK   100   100   001    -    515
170 Grown_Failing_Block_Ct  PO--CK   100   100   010    -    0
171 Program_Fail_Count      -O--CK   100   100   001    -    0
172 Erase_Fail_Count        -O--CK   100   100   001    -    0
173 Wear_Leveling_Count     PO--CK   100   100   010    -    0
174 Unexpect_Power_Loss_Ct  -O--CK   100   100   001    -    398
181 Non4k_Aligned_Access    -O---K   100   100   001    -    144 45 99
183 SATA_Iface_Downshift    -O--CK   100   100   001    -    0
184 End-to-End_Error        PO--CK   100   100   050    -    0
187 Reported_Uncorrect      -O--CK   100   100   001    -    0
188 Command_Timeout         -O--CK   100   100   001    -    0
189 Factory_Bad_Block_Ct    -OSR--   100   100   001    -    215
194 Temperature_Celsius     -O---K   100   100   000    -    0
195 Hardware_ECC_Recovered  -O-RCK   100   100   001    -    0
196 Reallocated_Event_Count -O--CK   100   100   001    -    0
197 Current_Pending_Sector  -O--CK   100   100   001    -    0
198 Offline_Uncorrectable   ----CK   100   100   001    -    0
199 UDMA_CRC_Error_Count    -O--CK   100   100   001    -    0
202 Perc_Rated_Life_Used    ---RC-   100   100   001    -    0
206 Write_Error_Rate        -OSR--   100   100   001    -    0
                            ||||||_ K auto-keep
                            |||||__ C event count
                            ||||___ R error rate
                            |||____ S speed/performance
                            ||_____ O updated online
                            |______ P prefailure warning

Read SMART Log Directory failed: scsi error medium or hardware error (serious)

General Purpose Log Directory Version 1
Address    Access  R/W   Size  Description
0x00       GPL     R/O      1  Log Directory
0x03       GPL     R/O  16383  Ext. Comprehensive SMART error log
0x04       GPL     R/O    255  Device Statistics log
0x07       GPL     R/O   3449  Extended self-test log
0x10       GPL     R/O      1  NCQ Command Error log
0x11       GPL     R/O      1  SATA Phy Event Counters
0x80-0x9f  GPL     R/W     16  Host vendor specific log
0xa0       GPL     VS    2000  Device vendor specific log
0xa1-0xbf  GPL     VS       1  Device vendor specific log
0xc0       GPL     VS      80  Device vendor specific log
0xc1-0xdf  GPL     VS       1  Device vendor specific log
0xe0       GPL     R/W      1  SCT Command/Status
0xe1       GPL     R/W      1  SCT Data Transfer

SMART Extended Comprehensive Error Log size 16383 not supported

Read SMART Error Log failed: scsi error medium or hardware error (serious)

SMART Extended Self-test Log size 3449 not supported

Read SMART Self-test Log failed: scsi error medium or hardware error (serious)

Read SMART Selective Self-test Log failed: scsi error medium or
hardware error (serious)

SCT Status Version:                  3
SCT Version (vendor specific):       1 (0x0001)
SCT Support Level:                   0
Device State:                        Stand-by (1)
Current Temperature:                  0 Celsius
Power Cycle Max Temperature:          0 Celsius
Lifetime    Max Temperature:          0 Celsius

SCT Temperature History Version:     2
Temperature Sampling Period:         10 minutes
Temperature Logging Interval:        10 minutes
Min/Max recommended Temperature:      0/70 Celsius
Min/Max Temperature Limit:           -5/75 Celsius
Temperature History Size (Index):    478 (38)

Index    Estimated Time   Temperature Celsius
  39    2014-08-13 02:20     ?  -
 ...    ..(475 skipped).    ..  -
  37    2014-08-16 09:40     ?  -
  38    2014-08-16 09:50     0  -

Write SCT (Get) Error Recovery Control Command failed: scsi error
medium or hardware error (serious)
SCT (Get) Error Recovery Control command failed

Device Statistics (GP Log 0x04)
Page Offset Size         Value  Description
  1  =====  =                =  == General Statistics (rev 2) ==
  1  0x008  4              515  Lifetime Power-On Resets
  1  0x010  4             4439  Power-on Hours
  1  0x018  6        933634265  Logical Sectors Written
  1  0x020  6          2708986  Number of Write Commands
  1  0x028  6        852431906  Logical Sectors Read
  1  0x030  6          6049196  Number of Read Commands
  4  =====  =                =  == General Errors Statistics (rev 1) ==
  4  0x008  4                0  Number of Reported Uncorrectable Errors
  4  0x010  4                0  Resets Between Cmd Acceptance and Completion
  5  =====  =                =  == Temperature Statistics (rev 1) ==
  5  0x008  1                0  Current Temperature
  5  0x010  1                0  Average Short Term Temperature
  5  0x018  1                0  Average Long Term Temperature
  5  0x020  1                0  Highest Temperature
  5  0x028  1                0  Lowest Temperature
  5  0x030  1                0  Highest Average Short Term Temperature
  5  0x038  1                0  Lowest Average Short Term Temperature
  5  0x040  1                0  Highest Average Long Term Temperature
  5  0x048  1                0  Lowest Average Long Term Temperature
  5  0x050  4                -  Time in Over-Temperature
  5  0x058  1               70  Specified Maximum Operating Temperature
  5  0x060  4                -  Time in Under-Temperature
  5  0x068  1                0  Specified Minimum Operating Temperature
  6  =====  =                =  == Transport Statistics (rev 1) ==
  6  0x008  4                0  Number of Hardware Resets
  6  0x010  4                0  Number of ASR Events
  6  0x018  4                0  Number of Interface CRC Errors
  7  =====  =                =  == Solid State Device Statistics (rev 1) ==
  7  0x008  1               21~ Percentage Used Endurance Indicator
                              |_ ~ normalized value

SATA Phy Event Counters (GP Log 0x11)
ID      Size     Value  Description
0x0001  4            0  Command failed due to ICRC error
0x000a  4            0  Device-to-host register FISes sent due to a COMRESET

On 14 August 2014 20:08, Mike <[email protected]> wrote:
> Hello again
>
> I've put the drive back into the system it was used in, and used the
> one other SATA port available.  Running smartctl returns some strange
> output, I've also included the kernel message in case it helps.
>
> root@voyage:~/smartmontools-6.3# ./smartctl -T permissive -x /dev/sda
> smartctl 6.3 2014-07-26 r3976 [i686-linux-3.4.4-voyage] (local build)
> Copyright (C) 2002-14, Bruce Allen, Christian Franke, www.smartmontools.org
>
> === START OF INFORMATION SECTION ===
> Vendor:               /3:0:0:0
> Product:
> Compliance:           SPC-5
> User Capacity:        600,332,565,813,390,450 bytes [600 PB]
> Logical block size:   774843950 bytes
> scsiModePageOffset: response length too short, resp_len=47 offset=50 bd_len=46
> scsiModePageOffset: response length too short, resp_len=47 offset=50 bd_len=46
>>> Terminate command early due to bad response to IEC mode page
> Read Cache is:        Unavailable
> Writeback Cache is:   Unavailable
>
> === START OF READ SMART DATA SECTION ===
> Log Sense failed, IE page [scsi response fails sanity test]
> Error Counter logging not supported
>
> scsiModePageOffset: response length too short, resp_len=47 offset=50 bd_len=46
> Device does not support Self Test logging
> Device does not support Background scan results logging
>
>
> [    3.899820] ata1: SATA max UDMA/133 abar m1024@0xffe40000 port
> 0xffe40100 irq 45
> [    3.899954] ata2: DUMMY
> [    3.900067] ata3: SATA max UDMA/133 abar m1024@0xffe40000 port
> 0xffe40200 irq 45
> [    3.900193] ata4: DUMMY
> [    4.205145] ata3: SATA link up 1.5 Gbps (SStatus 113 SControl 300)
> [    4.205279] ata1: SATA link down (SStatus 0 SControl 300)
> [    4.205686] ata3.00: failed to enable AA (error_mask=0x1)
> [    4.205772] ata3.00: ATA-9: M4-CT512M4SSD1, 070H, max UDMA/100
> [    4.205853] ata3.00: 1000215216 sectors, multi 16: LBA48 NCQ (depth 31/32)
> [    4.206476] ata3.00: failed to enable AA (error_mask=0x1)
> [    4.206568] ata3.00: configured for UDMA/100 (device error ignored)
> [    4.207143] scsi 3:0:0:0: Direct-Access     ATA      M4-CT512M4SSD1
>   070H PQ: 0 ANSI: 5
> [    4.207977] sd 3:0:0:0: [sda] 1000215216 512-byte logical blocks:
> (512 GB/476 GiB)
> [    4.208722] sd 3:0:0:0: [sda] Write Protect is off
> [    4.208822] sd 3:0:0:0: [sda] Mode Sense: 00 3a 00 00
> [    4.208998] sd 3:0:0:0: [sda] Write cache: enabled, read cache:
> enabled, doesn't support DPO or FUA
> [    4.569257] scsi 0:0:0:0: Direct-Access     TS       TS2GUFM-H
>   1100 PQ: 0 ANSI: 0 CCS
> [    4.571000] sd 0:0:0:0: [sdb] 4014080 512-byte logical blocks:
> (2.05 GB/1.91 GiB)
> [    4.572153] sd 0:0:0:0: [sdb] Write Protect is off
> [    4.572260] sd 0:0:0:0: [sdb] Mode Sense: 43 00 00 00
> [    4.573219] sd 0:0:0:0: [sdb] No Caching mode page present
> [    4.573323] sd 0:0:0:0: [sdb] Assuming drive cache: write through
> [    4.577728] sd 0:0:0:0: [sdb] No Caching mode page present
> [    4.577837] sd 0:0:0:0: [sdb] Assuming drive cache: write through
> [    4.579449]  sdb: sdb1
> [    4.583344] sd 0:0:0:0: [sdb] No Caching mode page present
> [    4.583454] sd 0:0:0:0: [sdb] Assuming drive cache: write through
> [    4.583561] sd 0:0:0:0: [sdb] Attached SCSI disk
> [    4.587875] sd 3:0:0:0: Attached scsi generic sg0 type 0
> [    4.588303] sd 0:0:0:0: Attached scsi generic sg1 type 0
> [   34.720249] ata3.00: exception Emask 0x0 SAct 0x1 SErr 0x0 action 0x6 frozen
> [   34.720368] ata3.00: failed command: READ FPDMA QUEUED
> [   34.720482] ata3.00: cmd 60/08:00:00:00:00/00:00:00:00:00/40 tag 0
> ncq 4096 in
> [   34.720489]          res 40/00:00:00:00:00/00:00:00:00:00/00 Emask
> 0x4 (timeout)
> [   34.720684] ata3.00: status: { DRDY }
> [   34.720769] ata3: hard resetting link
> [   40.074131] ata3: link is slow to respond, please be patient (ready=0)
> [   44.766167] ata3: COMRESET failed (errno=-16)
> [   44.766267] ata3: hard resetting link
> [   50.120136] ata3: link is slow to respond, please be patient (ready=0)
> [   54.812132] ata3: COMRESET failed (errno=-16)
> [   54.812216] ata3: hard resetting link
> [   60.166135] ata3: link is slow to respond, please be patient (ready=0)
> [   89.848141] ata3: COMRESET failed (errno=-16)
> [   89.848226] ata3: limiting SATA link speed to 1.5 Gbps
> [   89.848305] ata3: hard resetting link
> [   94.896060] ata3: COMRESET failed (errno=-16)
> [   94.896149] ata3: reset failed, giving up
> [   94.896225] ata3.00: disabled
> [   94.896300] ata3.00: device reported invalid CHS sector 0
> [   94.896402] ata3: EH complete
> [   94.896535] sd 3:0:0:0: [sda] Unhandled error code
> [   94.896613] sd 3:0:0:0: [sda]  Result: hostbyte=0x04 driverbyte=0x00
> [   94.896726] sd 3:0:0:0: [sda] CDB: cdb[0]=0x28: 28 00 00 00 00 00 00 00 08 00
> [   94.897243] end_request: I/O error, dev sda, sector 0
> [   94.897323] Buffer I/O error on device sda, logical block 0
> [   94.897517] sd 3:0:0:0: [sda] Unhandled error code
> [   94.897599] sd 3:0:0:0: [sda]  Result: hostbyte=0x04 driverbyte=0x00
> [   94.897712] sd 3:0:0:0: [sda] CDB: cdb[0]=0x28: 28 00 00 00 00 00 00 00 08 00
> [   94.898216] end_request: I/O error, dev sda, sector 0
> [   94.898293] Buffer I/O error on device sda, logical block 0
> [   94.898490] sd 3:0:0:0: [sda] Unhandled error code
> [   94.898568] sd 3:0:0:0: [sda]  Result: hostbyte=0x04 driverbyte=0x00
> [   94.898681] sd 3:0:0:0: [sda] CDB: cdb[0]=0x28: 28 00 00 00 00 00 00 00 08 00
> [   94.899194] end_request: I/O error, dev sda, sector 0
> [   94.899271] Buffer I/O error on device sda, logical block 0
> [   94.899383] ldm_validate_partition_table(): Disk read failed.
> [   94.899542] sd 3:0:0:0: [sda] Unhandled error code
> [   94.899620] sd 3:0:0:0: [sda]  Result: hostbyte=0x04 driverbyte=0x00
> [   94.899732] sd 3:0:0:0: [sda] CDB: cdb[0]=0x28: 28 00 00 00 00 00 00 00 08 00
> [   94.900237] end_request: I/O error, dev sda, sector 0
> [   94.900313] Buffer I/O error on device sda, logical block 0
> [   94.900506] sd 3:0:0:0: [sda] Unhandled error code
> [   94.900584] sd 3:0:0:0: [sda]  Result: hostbyte=0x04 driverbyte=0x00
> [   94.900697] sd 3:0:0:0: [sda] CDB: cdb[0]=0x28: 28 00 00 00 00 00 00 00 08 00
> [   94.901210] end_request: I/O error, dev sda, sector 0
> [   94.901287] Buffer I/O error on device sda, logical block 0
> [   94.901399]  sda: unable to read partition table
> [   94.902193] sd 3:0:0:0: [sda] READ CAPACITY(16) failed
> [   94.902281] sd 3:0:0:0: [sda]  Result: hostbyte=0x04 driverbyte=0x00
> [   94.902398] sd 3:0:0:0: [sda] Sense not available.
> [   94.902682] sd 3:0:0:0: [sda] READ CAPACITY failed
> [   94.902762] sd 3:0:0:0: [sda]  Result: hostbyte=0x04 driverbyte=0x00
> [   94.902877] sd 3:0:0:0: [sda] Sense not available.
> [   94.903338] sd 3:0:0:0: [sda] Got wrong page
> [   94.903419] sd 3:0:0:0: [sda] Assuming drive cache: write through
> [   94.903501] sd 3:0:0:0: [sda] Attached SCSI disk
>
> On 13 August 2014 22:27, Volker Kuhlmann <[email protected]> wrote:
>> On Wed 13 Aug 2014 03:36:23 NZST +1200, Mike wrote:
>>
>>> My m4 512Gb CT512M4SSD1 suddenly failed to read. Below is the smartctl
>>> output that returns eventually, taking at least a couple of minutes.  I've
>>> tried 2 different USB caddys and also SATA connection to laptop running
>>> livecd.  The drive has exclusively been mounted read-only on a Linux system
>>> with an ext4 filesystem, and has had unexpected power loss routinely
>>> throughout its life.  The values in the attribute table look OK to me I
>>> think, but the serious errors at the bottom look bad.  Any ideas greatly
>>> appreciated:
>>
>>> Read SMART Log Directory failed.
>>>
>>> Error SMART Error Log Read failed: scsi error medium or hardware error
>>> (serious)
>>> Smartctl: SMART Error Log Read Failed
>>> Error SMART Error Self-Test Log Read failed: scsi error medium or hardware
>>> error (serious)
>>> Smartctl: SMART Self Test Log Read Failed
>>> Error SMART Read Selective Self-Test Log failed: scsi error medium or
>>> hardware error (serious)
>>> Smartctl: SMART Selective Self Test Log Read Failed
>>
>> This looks like smartctl may not be able to communicate with the drive
>> properly, which is what you expect if the drive/computer interface
>> circuitry is damaged. However, these errors don't seem to exclude the
>> possibility of a serious platter failure - or rather the equivalent in
>> solid state.
>>
>> Moving the disk around between enclosures etc can easily damage the
>> drive electronics if sufficient ESD protection procedures have not been
>> followed (those procedures must be sufficient regardless of your opinion
>> of them ;-) ). If the problem first occurred without having moved any
>> hardware around then the problem is with the drive. If you can reliably
>> exclude the problem being located at the computer, enclosure or
>> connections/cables then it's also the drive.
>>
>> If the problem is with the drive then the drive is most certainly
>> finished. Copy off all the data you can asap and use your backups for
>> the rest.
>>
>> It might be a good idea to treat SSDs the same as spinning versions. The
>> differences are in speed and behaviour after being dropped, not
>> necessarily in long-term reliability of data.
>>
>> HTH,
>>
>> Volker
>>
>> --
>> Volker Kuhlmann
>> http://volker.top.geek.nz/      Please do not CC list postings to me.
>>
>> ------------------------------------------------------------------------------
>> _______________________________________________
>> Smartmontools-support mailing list
>> [email protected]
>> https://lists.sourceforge.net/lists/listinfo/smartmontools-support

------------------------------------------------------------------------------
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.