Re: [PATCH] powerpc/rtas_pci: No hotplug on permanently removed device on pSeries
Harsh Prateek Bora <[email protected]>
| Newsgroups | gmane.linux.ports.ppc.embedded |
|---|---|
| Message-ID | <c6117753-10a7-4e10-bd8f-526a5dd8a300__42850.5702687957$1787671015$gmane$org@linux.ibm.com> |
Hi Tasmiya, There was a v2 posted here: https://lore.kernel.org/linuxppc-dev/[email protected]/ Although a minor change in return value and re-ordering before CONFIG_BLOCKED incorporated in v2, it's better to test v2 and provide test feedback on the same. Thanks Harsh On 25/08/26 3:52 pm, tasmiya wrote: > Greetings, > > Tested-by: Tasmiya Nalatwad <[email protected]> > Reported-by: Tasmiya Nalatwad <[email protected]> > > > I have Tested the patch in my KVM guest environment with device > passthrough via vfio, It works fine and fixes the issue. > Triggered EEH freeze beyond eeh_max_freezes to permanently remove the > device, then hotplugged a new PCI device. Without the patch, the rescan > attempted to bring back the permanently removed device. With the patch > applied, the removed device is correctly skipped during rescan and only > newly hotplugged device is seen in the guest. > > > Thank you, > > Tasmiya > > > On 27/04/26 8:25 am, Shivaprasad G Bhat wrote: >> The eeh_driver disables and offlines the PE permanently when it >> exceeds the freeze count beyond eeh_max_freeze within the last hour. >> The PE is only offline, so the device tree entries, eeh device >> references are all intact till the real unplug of the device from >> the guest/host takes place. >> >> On pSeries, with a new hotplug of any PCI device, the drmgr initiates >> a system-wide PCI rescan, which finds devices offlined by the eeh_driver >> and there will be attempts to bring them online. This leads to >> recurring EEHs either at the config read time itself or a bit >> later depending on the type of the problem. >> >> For PowerNV, the commit d2b0f6f77ee5 ("powerpc/eeh: No hotplug on >> permanently removed dev") introduced the EEH_DEV_REMOVED flag to >> prevent such inadvertent rescans on hierarchical toplogies relavent in >> Baremetal setups. For pSeries, such topologies don't really make sense >> as the devices are either part of the same PE OR exposed as independent >> devices on multiple virtual PHBs. However, the inadvertent rescans are >> still a possibility with either hotplug of a new device or otherwise >> with manual system-wide pci bus rescan attempts. >> >> So the patch checks for EEH_DEV_REMOVED before allowing config space >> access just like PowerNV, making the PCI core omit the PE, and thus >> preventing subsequent EEH recurances. The patch is tested on PowerVM >> and KVM machines with single and multi-function devices, and on the >> devices behind a switch. The unplug of the affected devices post EEH >> removal is also working fine as expected. >> >> Signed-off-by: Shivaprasad G Bhat <[email protected]> >> References: d2b0f6f77ee5 ("powerpc/eeh: No hotplug on permanently >> removed dev") >> --- >> arch/powerpc/kernel/rtas_pci.c | 6 ++++++ >> 1 file changed, 6 insertions(+) >> >> diff --git a/arch/powerpc/kernel/rtas_pci.c b/arch/powerpc/kernel/ >> rtas_pci.c >> index fccf96e897f6..ce24b18712ca 100644 >> --- a/arch/powerpc/kernel/rtas_pci.c >> +++ b/arch/powerpc/kernel/rtas_pci.c >> @@ -57,6 +57,9 @@ int rtas_pci_dn_read_config(struct pci_dn *pdn, int >> where, int size, u32 *val) >> if (pdn->edev && pdn->edev->pe && >> (pdn->edev->pe->state & EEH_PE_CFG_BLOCKED)) >> return PCIBIOS_SET_FAILED; >> + >> + if (pdn->edev && pdn->edev->mode & EEH_DEV_REMOVED) >> + return PCIBIOS_SET_FAILED; >> #endif >> addr = rtas_config_addr(pdn->busno, pdn->devfn, where); >> @@ -108,6 +111,9 @@ int rtas_pci_dn_write_config(struct pci_dn *pdn, >> int where, int size, u32 val) >> if (pdn->edev && pdn->edev->pe && >> (pdn->edev->pe->state & EEH_PE_CFG_BLOCKED)) >> return PCIBIOS_SET_FAILED; >> + >> + if (pdn->edev && pdn->edev->mode & EEH_DEV_REMOVED) >> + return PCIBIOS_SET_FAILED; >> #endif >> addr = rtas_config_addr(pdn->busno, pdn->devfn, where); >> >> >