Re: [PATCH v3] cxl/mce: Avoid alias page retirement for corrected errors

Alison Schofield <[email protected]>
Newsgroups org.kernel.vger.stable,org.kernel.vger.linux-cxl,org.kernel.vger.linux-edac,org.kernel.vger.linux-kernel
Message-ID <[email protected]>
On Mon, Aug 24, 2026 at 07:19:57PM +0530, Shaikh Kamaluddin wrote:
> cxl_handle_mce() offlines the aliased page of an Extended Linear Cache
> (ELC) region for any MCE with a usable address in the region. This
> includes corrected errors, needlessly reducing usable memory.
> 
> Skip ELC alias retirement when mce_is_correctable() identifies the
> reported error as corrected. Uncorrected errors with a usable address
> in the ELC region continue to retire the aliased page as before.
> 
> mce_usable_address() already implements the per-vendor checks that
> determine whether the reported address is usable, so no separate
> memory-error classification is needed.
> 
> On AMD, a corrected legacy bank 4 DRAM ECC error (XEC 8) reports a
> usable address and retired the aliased page before this change.
> Tested under QEMU with Intel Skylake-Server and AMD EPYC-Milan models.
> 
> Fixes: 516e5bd0b6bf ("cxl: Add mce notifier to emit aliased address for extended linear cache")
> Signed-off-by: Shaikh Kamaluddin <[email protected]>

Reviewed-by: Alison Schofield <[email protected]>

To be complete, I'm not retesting since v2, so not offering my
testing tag here. I have confidence in your thorough testing!

-- Alison



> ---
> Changelog:
> 
> v2 -> v3:
>  - Drop the mce_is_memory_error() gate added in v2. On AMD it rejects
>    every non-UMC bank type, so poison-consumption records reported from
>    Load/Store or Data Fabric banks never reached mce_usable_address(),
>    which would have accepted them via MCA_STATUS[Poison]. (Sashiko)
>  - mce_usable_address() already implements the complete per-vendor
>    address-usability logic, so no separate memory-error classification
>    is needed.
>  - Add AMD test coverage (QEMU/TCG, EPYC-Milan) alongside the existing
>    Intel results.
> 
> Link to v2: https://lore.kernel.org/linux-cxl/[email protected]/
> 
> v1 -> v2:
>   - Reworked commit message per review feedback
>   - Dropped redundant comment from the code
> 
> Link to v1: https://lore.kernel.org/linux-cxl/[email protected]/ 
> 
> Reproduced under QEMU/vng. QEMU's cxl-type3 does not emulate the HMAT
> extended-linear address_mode bit, so ELC was forced locally for testing
> via a one-line debug hack in cxl_region_probe() (not part of this patch):
> 
> 	if (!p->cache_size && p->res)
> 		p->cache_size = resource_size(p->res) / 2;
> 
> Steps:
> 1. Boot with an Intel CPU model under TCG (KVM host-passthrough will
>     otherwise leak the host's real vendor ID, and AMD/SMCA takes a
>     different mce_usable_address() path, covered in the section below):
> 
> vng -v -r ./arch/x86/boot/bzImage --disable-kvm --qemu-opts='-cpu Skylake-Server-v4,+mce,+mca -m 4G -machine q35,cxl=on -object memory-backend-ram,id=cxl-mem0,size=512M -device pxb-cxl,bus_nr=12,bus=pcie.0,id=cxl.0 -device cxl-rp,port=0,bus=cxl.0,id=root_port0,chassis=0,slot=0 -device cxl-type3,bus=root_port0,volatile-memdev=cxl-mem0,id=cxl-mem-device0 -M cxl-fmw.0.targets.0=cxl.0,cxl-fmw.0.size=512M'
> 
> 2. modprobe mce-inject
> $ cxl list -M
> $ cxl list -D
> $ cxl create-region -m mem0 -d decoder0.0 -w 1 -g 256 -t ram
> $ dmesg | grep "DEBUG: forced cache_size"
> $ modprobe device_dax
> $ modprobe kmem
> $ ls /sys/bus/dax/devices/
> $ daxctl reconfigure-device dax0.0 --mode=system-ram
> $ lsmem
> RANGE                                  SIZE  STATE REMOVABLE BLOCK
> 0x0000000000000000-0x000000007fffffff    2G online       yes  0-15
> 0x0000000100000000-0x000000017fffffff    2G online       yes 32-47
> 0x0000000190000000-0x00000001afffffff  512M online       yes 50-53
> 
> Memory block size:       128M
> Total online memory:     4.5G
> Total offline memory:      0B
> 
> 3. Load the injector if not loaded earlier and derive the two MCi_STATUS values.
> 
>    # modprobe mce-inject
> 
>     MCi_STATUS bit layout used here (arch/x86/include/asm/mce.h):
>       bit 63  VAL     - record valid
>       bit 61  UC      - uncorrected (0 = corrected error under test)
>       bit 60  EN      - error reporting enabled
>       bit 59  MISCV   - MCi_MISC valid
>       bit 58  ADDRV   - MCi_ADDR valid
>       bits[15:0] MCACOD - retained from the v2 test setup so the Intel
>                           and AMD runs differ only in CPU model; not
>                           consulted by this path.
> 
>     python3 -c "
>     VAL, UC, EN, MISCV, ADDRV = 1<<63, 1<<61, 1<<60, 1<<59, 1<<58
>     MCACOD_MEM = 1<<7                
>     ce = VAL | EN | MISCV | ADDRV | MCACOD_MEM
>     uc = ce | UC
>     print(f'CE status = {hex(ce)}')
>     print(f'UC status = {hex(uc)}')"
>     # CE status = 0x9c00000000000080
>     # UC status = 0xbc00000000000080
> 
>     MCi_MISC: address-mode field, bits[8:6], must be 2 (physical):
> 
>     python3 -c "print(hex(2 << 6))"
>     # misc = 0x80
> 
> 4.  # cd /sys/kernel/debug/mce-inject
> # echo sw > flags
> # echo 0x9C00000000000080 > status
> # echo 0x80 > misc
> # echo 0x190010000 > addr
> # echo 9 > bank
> mce: [Hardware Error]: Machine check events logged
> cxl_mce_debug: entered status=0x9c00000000000080 addr=0x190010000 cache_size=0x10000000 res=[mem 0x190000000-0x1afffffff flags 0x200] usable=1
> cxl_mce_debug: spa=0x190010000 contains=1
> cxl_mce_debug: spa_alias=0x1a0010000 pfn=0x1a0010 pfn_valid=1
> cxl_region region0: Offlining aliased SPA address0: 0x1a0010000
> Memory failure: 0x1a0010: recovery action for free buddy page: Recovered
> mce: [Hardware Error]: CPU 0: Machine Check: 0 Bank 9: 9c00000000000080
> mce: [Hardware Error]: TSC a605ccda40 ADDR 190010000 MISC 80 
> mce: [Hardware Error]: PROCESSOR 0:50654 TIME 1786375908 SOCKET 0 APIC 0 microcode 1
> # grep HardwareCorrupted /proc/meminfo
> HardwareCorrupted:     4 kB
> ---------------------------------------
> 
> CE without this patch:
>   cxl_region region0: Offlining aliased SPA address0: 0x1a0010000
>   Memory failure: 0x1a0010: recovery action for free buddy page: Recovered
>   HardwareCorrupted: 4 kB
> ------------------------------------- 
> 
> CE with this patch:
>   (no "Offlining aliased SPA" message logged)
>   HardwareCorrupted: 0 kB
> 
> # cd /sys/kernel/debug/mce-inject
> # echo sw > flags
> # echo 0x9C00000000000080 > status
> # echo 0x80 > misc
> # echo 0x190010000 > addr
> # echo 9 > bank
> mce: [Hardware Error]: Machine check events logged
> cxl_mce_debug: entered status=0x9c00000000000080 addr=0x190010000 cache_size=0x10000000 res=[mem 0x190000000-0x1afffffff flags 0x200] usable=1
> mce: [Hardware Error]: CPU 0: Machine Check: 0 Bank 9: 9c00000000000080
> mce: [Hardware Error]: TSC 2eeae32fe0 ADDR 190010000 MISC 80 
> mce: [Hardware Error]: PROCESSOR 0:50654 TIME 1786377156 SOCKET 0 APIC 0 microcode 1
> clocksource: Watchdog remote CPU 11 read timed out
> 
> # grep HardwareCorrupted /proc/meminfo
> HardwareCorrupted:     0 kB
> --------------------------------------
> 
> For UC, same steps only status bit information will change :
> 
> # cd /sys/kernel/debug/mce-inject
> # echo sw > flags
> # echo 0xbc00000000000080 > status
> # echo 0x80 > misc
> # echo 0x190010000 > addr
> # echo 9 > bank
> # dmesg | grep cxl_mce_debug
> #  grep HardwareCorrupted /proc/meminfo
> mce: [Hardware Error]: Machine check events logged
> cxl_mce_debug: entered status=0xbc00000000000080 addr=0x190010000 cache_size=0x10000000 res=[mem 0x190000000-0x1afffffff flags 0x200] usable=1
> cxl_mce_debug: spa=0x190010000 contains=1
> cxl_mce_debug: spa_alias=0x1a0010000 pfn=0x1a0010 pfn_valid=1
> cxl_region region0: Offlining aliased SPA address0: 0x1a0010000
> Memory failure: 0x1a0010: recovery action for free buddy page: Recovered
> mce: [Hardware Error]: CPU 0: Machine Check: 0 Bank 9: bc00000000000080
> mce: [Hardware Error]: TSC e9dd2869e0 ADDR 190010000 MISC 80 
> mce: [Hardware Error]: PROCESSOR 0:50654 TIME 1786377750 SOCKET 0 APIC 0 microcode 1
> 
> Patched, UC: alias still offlined, confirming uncorrected handling is unchanged by this patch.
>     cxl_region region0: Offlining aliased SPA address0: 0x1a0010000
>     Memory failure: 0x1a0010: recovery action for free buddy page: Recovered
>     HardwareCorrupted: 4 kB
> -----------------------------------------------------------------------
> 
> AMD Platform Testing:
> ---------------------
> 
> Note the vendor difference: on Intel, MISCV with MCi_MISC[8:6]=2 makes
> any ADDRV record usable, so a plain CE is address-usable. On AMD an
> address is usable only via poison, legacy bank-4 DRAM ECC, or PADDRV --
> hence the differing usable= values for the same status word below.
> 
> 
> vendor_id	: AuthenticAMD
> model name	: AMD EPYC-Milan Processor
> 
> -------------------------------------
> Mainline Kernel Without this patch:
> -------------------------------------
> 
> vng -v -r ./arch/x86/boot/bzImage --disable-kvm --qemu-opts='-cpu EPYC-Milan -m 4G -machine q35,cxl=on -object memory-backend-ram,id=cxl-mem0,size=512M -device pxb-cxl,bus_nr=12,bus=pcie.0,id=cxl.0 -device cxl-rp,port=0,bus=cxl.0,id=root_port0,chassis=0,slot=0 -device cxl-type3,bus=root_port0,volatile-memdev=cxl-mem0,id=cxl-mem-device0 -M cxl-fmw.0.targets.0=cxl.0,cxl-fmw.0.size=512M'
> 
> # cxl list -M
> # cxl list -D
> # cxl create-region -m mem0 -d decoder0.0 -w 1 -g 256 -t ram
> # dmesg | grep "DEBUG: forced cache_size"
> # modprobe device_dax
> # modprobe kmem
> # daxctl reconfigure-device dax0.0 --mode=system-ram
> 
> # lsmem
> RANGE                                  SIZE  STATE REMOVABLE BLOCK
> 0x0000000000000000-0x000000007fffffff    2G online       yes  0-15
> 0x0000000100000000-0x000000017fffffff    2G online       yes 32-47
> 0x0000000190000000-0x00000001afffffff  512M online       yes 50-53
> 
> Memory block size:       128M
> Total online memory:     4.5G
> Total offline memory:      0B
> Load Injector
> # modprobe mce-inject
> --------------------------------------------------------------------------------
> 
> case1:
> 
> label              status           addr         bank   note
> 
> [LEGACY_B4_CE]  0x9c00000000080000  0x190070000   4     "case 3: bank4 XEC8 corrected+usable "[Before Patch]
> ---------------------------------------------------------------------------------------------------------------------
> 
> # cd /sys/kernel/debug/mce-inject
> # echo sw > flags
> # echo 0x9c00000000080000 > status
> # echo 0x80 > misc
> # echo 0x190070000 > addr
> # echo 4 > bank
> mce: [Hardware Error]: Machine check events logged
> cxl_mce_debug: entered status=0x9c00000000080000 addr=0x190070000 cache_size=0x10000000 res=[mem 0x190000000-0x1afffffff flags 0x200] usable=1
> cxl_mce_debug: spa=0x190070000 contains=1
> cxl_mce_debug: spa_alias=0x1a0070000 pfn=0x1a0070 pfn_valid=1
> cxl_region region0: Offlining aliased SPA address0: 0x1a0070000
> Memory failure: 0x1a0070: recovery action for free buddy page: Recovered
> mce: [Hardware Error]: CPU 0: Machine Check: 0 Bank 4: 9c00000000080000
> mce: [Hardware Error]: TSC cbdf882d20 ADDR 190070000 MISC 80 
> mce: [Hardware Error]: PROCESSOR 2:a00f11 TIME 1787479138 SOCKET 0 APIC 0 microcode 1000065
> # grep HardwareCorrupted /proc/meminfo
> HardwareCorrupted:     4 kB
> --------------------------------------------------------------------------------------------
> 
> case2:
> label        status              addr         bank   note
> 
> [POISON]  0x9c00080000000080    0x190040000    9     "case 2: poison only (UC=0) - divergence to note"[Before Patch]
> -----------------------------------------------------------------------------------------------------------------------
> 
> # echo sw > flags
> # echo 0x9c00080000000080 > status
> # echo 0x80 > misc
> # echo 0x190040000 > addr
> # echo 9 > bank
> mce: [Hardware Error]: Machine check events logged
> cxl_mce_debug: entered status=0x9c00080000000080 addr=0x190040000 cache_size=0x10000000 res=[mem 0x190000000-0x1afffffff flags 0x200] usable=1
> cxl_mce_debug: spa=0x190040000 contains=1
> cxl_mce_debug: spa_alias=0x1a0040000 pfn=0x1a0040 pfn_valid=1
> cxl_region region0: Offlining aliased SPA address0: 0x1a0040000
> Memory failure: 0x1a0040: recovery action for free buddy page: Recovered
> mce: [Hardware Error]: CPU 0: Machine Check: 0 Bank 9: 9c00080000000080
> mce: [Hardware Error]: TSC 17153949200 ADDR 190040000 MISC 80 
> mce: [Hardware Error]: PROCESSOR 2:a00f11 TIME 1787479277 SOCKET 0 APIC 0 microcode 1000065
> # grep HardwareCorrupted /proc/meminfo
> HardwareCorrupted:     8 kB (cumulative added previous 4kb + 4kb this test)
> -----------------------------------------------
> 
> AMD Platform Testing with applied patch:
> ------------------------------------------
> 
> # cxl list -D
> # cxl create-region -m mem0 -d decoder0.0 -w 1 -g 256 -t ram
> # dmesg | grep "DEBUG: forced cache_size"
> # modprobe device_dax
> # modprobe kmem
> # daxctl reconfigure-device dax0.0 --mode=system-ram
> # lsmem
> # modprobe mce-inject
> ----------------------------------------------------
> 
> case1:
> label              status           addr         bank   note
> 
> [ CE]          0x9c00000000000080    0x190010000   9     "case 1: corrected, no usable addr"
> ----------------------------------------------------------------------------------------
> 
> # cd /sys/kernel/debug/mce-inject
> # echo sw > flags
> # echo 0x9c00000000000080 > status
> # echo 0x80 > misc
> # echo 0x190010000 > addr
> # echo 9 > bank
> mce: [Hardware Error]: Machine check events logged
> cxl_mce_debug: entered status=0x9c00000000000080 addr=0x190010000 cache_size=0x10000000 res=[mem 0x190000000-0x1afffffff flags 0x200] usable=0
> mce: [Hardware Error]: CPU 0: Machine Check: 0 Bank 9: 9c00000000000080
> mce: [Hardware Error]: TSC 8e59f2f3bc0 ADDR 190010000 MISC 80 
> mce: [Hardware Error]: PROCESSOR 2:a00f11 TIME 1787454918 SOCKET 0 APIC 0 microcode 1000065
> # grep HardwareCorrupted /proc/meminfo
> HardwareCorrupted:     0 kB
> -------------------------------------------------------------------------
> 
> case2:
> label              status           addr         bank   note
> [ UC ]        0xbc00000000000080    0x190020000   9      "case 1: uncorrected, no poison -> unusable"
> -------------------------------------------------------------------------------------------------
> 
> cd /sys/kernel/debug/mce-inject
> echo sw > flags
> echo 0xbc00000000000080 > status
> echo 0x80 > misc
> echo 0x190020000 > addr
> echo 9 > bank
> Result:
> mce: [Hardware Error]: Machine check events logged
> cxl_mce_debug: entered status=0xbc00000000000080 addr=0x190020000 cache_size=0x10000000 res=[mem 0x190000000-0x1afffffff flags 0x200] usable=0
> mce: [Hardware Error]: CPU 0: Machine Check: 0 Bank 9: bc00000000000080
> mce: [Hardware Error]: TSC 134d6de13240 ADDR 190020000 MISC 80 
> mce: [Hardware Error]: PROCESSOR 2:a00f11 TIME 1787460569 SOCKET 0 APIC 0 microcode 1000065
> # grep HardwareCorrupted /proc/meminfo
> HardwareCorrupted:     0 kB
> -------------------------------------------------------------------
> 
> case3:
> label              status           addr         bank   note
> 
> [ POISON ]   0x9c00080000000080    0x190040000   9      "case 2: poison only (UC=0) - divergence to note"
> ------------------------------------------------------------------------------------------------------------
> 
> # cd /sys/kernel/debug/mce-inject
> # echo sw > flags
> # echo 0x9c00080000000080 > status
> # echo 0x80 > misc
> # echo 0x190040000 > addr
> # echo 9 > bank
> Result:
> mce: [Hardware Error]: Machine check events logged
> cxl_mce_debug: entered status=0x9c00080000000080 addr=0x190040000 cache_size=0x10000000 res=[mem 0x190000000-0x1afffffff flags 0x200] usable=1
> mce: [Hardware Error]: CPU 0: Machine Check: 0 Bank 9: 9c00080000000080
> mce: [Hardware Error]: TSC 14e29da617c0 ADDR 190040000 MISC 80 
> mce: [Hardware Error]: PROCESSOR 2:a00f11 TIME 1787463321 SOCKET 0 APIC 0 microcode 1000065
> # grep HardwareCorrupted /proc/meminfo
> HardwareCorrupted:     0 kB
> ------------------------------------------
> 
> case4:
> label              status           addr         bank   note
> 
> [ DEFERRED ]  0x9c00100000000080    0x190020000   9     "case 1: deferred, no poison -> unusable"
> ------------------------------------------------------------------------------------------------
> 
> # echo sw > flags
> # echo 0x9c00100000000080 > status
> # echo 0x80 > misc
> #echo 0x190020000 > addr
> #echo 9 > bank
> Result:
> mce: [Hardware Error]: Machine check events logged
> cxl_mce_debug: entered status=0x9c00100000000080 addr=0x190020000 cache_size=0x10000000 res=[mem 0x190000000-0x1afffffff flags 0x200] usable=0
> mce: [Hardware Error]: CPU 0: Machine Check: 0 Bank 9: 9c00100000000080
> mce: [Hardware Error]: TSC 9a2d801ff40 ADDR 190020000 
> mce: [Hardware Error]: PROCESSOR 2:a00f11 TIME 1787456818 SOCKET 0 APIC 0 microcode 1000065
> # grep HardwareCorrupted /proc/meminfo
> HardwareCorrupted:     0 kB
> -----------------------------------------------
> 
> case5:
> label              status            addr         bank    note
> 
> [ DEF_POISON ]  0x9c00180000000080    0x190060000   9      "case 2: deferred+poison - MUST be handled"
> --------------------------------------------------------------------------------------------------------
> 
> # echo sw > flags
> # echo 0x9c00180000000080 > status
> # echo 0x80 > misc
> # echo 0x190060000 > addr
> # echo 9 > bank
> Result:
> mce: [Hardware Error]: Machine check events logged
> cxl_mce_debug: entered status=0x9c00180000000080 addr=0x190060000 cache_size=0x10000000 res=[mem 0x190000000-0x1afffffff flags 0x200] usable=1
> cxl_mce_debug: spa=0x190060000 contains=1
> cxl_mce_debug: spa_alias=0x1a0060000 pfn=0x1a0060 pfn_valid=1
> cxl_region region0: Offlining aliased SPA address0: 0x1a0060000
> Memory failure: 0x1a0060: recovery action for free buddy page: Recovered
> mce: [Hardware Error]: CPU 0: Machine Check: 0 Bank 9: 9c00180000000080
> mce: [Hardware Error]: TSC a60fd45ae60 ADDR 190060000 MISC 80 
> mce: [Hardware Error]: PROCESSOR 2:a00f11 TIME 1787457072 SOCKET 0 APIC 0 microcode 1000065
> # grep HardwareCorrupted /proc/meminfo
> HardwareCorrupted:     4 kB
> -----------------------------------
> 
> case6:
> label              status           addr         bank   note
> 
> [UC_POISON]  0xbc00080000000080    0x190050000   9       "case 2: poison consumption - MUST be handled"
> --------------------------------------------------------------------------------------------------------
> 
> # echo sw > flags
> # echo 0xbc00080000000080 > status
> # echo 0x80 > misc
> # echo 0x190050000 > addr
> # echo 9 > bank
> 
> Resul:
> 
> mce: [Hardware Error]: Machine check events logged
> cxl_mce_debug: entered status=0xbc00080000000080 addr=0x190050000 cache_size=0x10000000 res=[mem 0x190000000-0x1afffffff flags 0x200] usable=1
> cxl_mce_debug: spa=0x190050000 contains=1
> cxl_mce_debug: spa_alias=0x1a0050000 pfn=0x1a0050 pfn_valid=1
> cxl_region region0: Offlining aliased SPA address0: 0x1a0050000
> Memory failure: 0x1a0050: recovery action for free buddy page: Recovered
> mce: [Hardware Error]: CPU 0: Machine Check: 0 Bank 9: bc00080000000080
> mce: [Hardware Error]: TSC f03abeed540 ADDR 190050000 MISC 80 
> mce: [Hardware Error]: PROCESSOR 2:a00f11 TIME 1787457328 SOCKET 0 APIC 0 microcode 1000065
> # grep HardwareCorrupted /proc/meminfo
> HardwareCorrupted:     4 kB
> -------------------------------------------------------------------------------------
> 
> case7:
> label              status                   addr         bank    note
> 
> [B4_NONMEM_POISON]   0x9c00080000000080    0x190080000   4       "case 3 override: bank4 non-mem, poison ignored"
> -------------------------------------------------------------------------------------------------------------------
> 
> # echo sw > flags
> # echo 0x9c00080000000080  > status
> # echo 0x80 > misc
> # echo 0x190080000 > addr
> # echo 4 > bank
> 
> Result:
> 
> mce: [Hardware Error]: Machine check events logged
> cxl_mce_debug: entered status=0x9c00080000000080 addr=0x190080000 cache_size=0x10000000 res=[mem 0x190000000-0x1afffffff flags 0x200] usable=0
> mce: [Hardware Error]: CPU 0: Machine Check: 0 Bank 4: 9c00080000000080
> mce: [Hardware Error]: TSC 21c66cecf80 ADDR 190080000 MISC 80 
> mce: [Hardware Error]: PROCESSOR 2:a00f11 TIME 1787454082 SOCKET 0 APIC 0 microcode 1000065
> # grep HardwareCorrupted /proc/meminfo
> HardwareCorrupted:     0 kB
> --------------------------------------------------------------------------------------------
> 
> case8:
> label              status               addr         bank     note
> 
> [LEGACY_B4_CE]   0x9c00000000080000    0x190070000   4        "case 3: bank4 XEC8 corrected+usable - THE FIX"
> -------------------------------------------------------------------------------------------------------------------
> 
> echo sw > flags
> echo 0x9c00000000080000 > status
> echo 0x80 > misc
> echo 0x190070000 > addr
> echo 4 > bank
> 
> result:
> -------
> 
> mce: [Hardware Error]: Machine check events logged
> cxl_mce_debug: entered status=0x9c00000000080000 addr=0x190070000 cache_size=0x10000000 res=[mem 0x190000000-0x1afffffff flags 0x200] usable=1
> mce: [Hardware Error]: CPU 0: Machine Check: 0 Bank 4: 9c00000000080000
> mce: [Hardware Error]: TSC 3610305ed80 ADDR 190070000 MISC 80 
> mce: [Hardware Error]: PROCESSOR 2:a00f11 TIME 1787454482 SOCKET 0 APIC 0 microcode 1000065
> # grep HardwareCorrupted /proc/meminfo
> HardwareCorrupted:     0 kB
> -------------------------------------------
> 
> Note: an AMD record with MCI_STATUS_POISON set but UC and Deferred
> clear is filtered by this patch where it previously was not. Per AMD64
> APM 24593 Rev 3.45, when UC is clear the error class is determined
> solely by the Deferred bit, and Poison qualifies an uncorrected error
> rather than establishing one. Real poison is therefore always reported
> with Deferred set (not yet consumed) or UC set (consumed via #MC), and
> both of those cases still retire the alias. Poison with UC and Deferred
> both clear is not a valid encoding and appears only under software
> injection.
> 
>  drivers/cxl/core/mce.c | 8 +++++++-
>  1 file changed, 7 insertions(+), 1 deletion(-)
> 
> diff --git a/drivers/cxl/core/mce.c b/drivers/cxl/core/mce.c
> index 65fed913b221..3ac6802e750d 100644
> --- a/drivers/cxl/core/mce.c
> +++ b/drivers/cxl/core/mce.c
> @@ -18,7 +18,13 @@ static int cxl_handle_mce(struct notifier_block *nb, unsigned long val,
>  	u64 spa, spa_alias;
>  	unsigned long pfn;
>  
> -	if (!mce || !mce_usable_address(mce))
> +	if (!mce)
> +		return NOTIFY_DONE;
> +
> +	if (mce_is_correctable(mce))
> +		return NOTIFY_DONE;
> +
> +	if (!mce_usable_address(mce))
>  		return NOTIFY_DONE;
>  
>  	spa = mce->addr & MCI_ADDR_PHYSADDR;
> 
> base-commit: 7098e9cd98a05c0c5de2fae0c2465f9d966fdd07
> -- 
> 2.43.0
>
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.