Re: [PATCH] PCI: Drop unnecessary retries when restoring BARs
"Rafael J. Wysocki" <[email protected]>
| Newsgroups | dev.linux.lists.sashiko,org.kernel.vger.linux-pci |
|---|---|
| Message-ID | <CAJZ5v0g3f8_7v-8W0SAbn_8vmKUJgpDbgyYMSjv03QjXJ8bCHQ@mail.gmail.com> |
On Mon, May 4, 2026 at 7:09 PM Bjorn Helgaas <[email protected]> wrote: > > [+cc Mario] > > On Mon, May 04, 2026 at 09:49:36AM +0200, Lukas Wunner wrote: > > On Sun, May 03, 2026 at 01:51:08PM +0000, [email protected] wrote: > > > Thank you for your contribution! Sashiko AI review found 1 potential > > > issue(s) to consider: > > > - [High] Removing the read-back and retry loop for BAR restoration in > > > `pci_restore_state()` introduces a risk of silent regressions for > > > hardware resuming from non-FLR resets (such as D3hot to D0 transitions > > > or custom driver resets). The commit incorrectly assumes the 60s delay > > > from `pci_dev_wait()` covers all usages, but standard PM resume paths > > > only delay for 10ms (`PCI_PM_D3_WAIT`) before calling `pci_restore_state()`. > > > Historically, hardware that needed slightly longer to accept > > > configuration writes relied on the 10x 1ms retry loop to successfully > > > restore BARs. By removing both the retry and the read-back verification, > > > BAR writes to slow devices will be silently dropped, leaving hardware > > > unconfigured and causing MMIO accesses to result in IOMMU faults or > > > kernel crashes. > > > > Hallucination alert: > > > > PCI_PM_D3_WAIT does not exist, it was renamed to PCI_PM_D3HOT_WAIT > > six years ago by commit 3789af9a13e5, which went into v5.10. > > > > The macro is used in: > > > > pci_pm_resume_noirq() > > pci_pm_default_resume_early() > > pci_pm_power_up_and_verify_state() > > pci_power_up() > > pci_dev_d3_sleep() > > > > However before pci_power_up() calls pci_dev_d3_sleep(), it reads > > the PMCSR register and errors out if config space is inaccessible. > > > > Hence when pci_restore_state() is invoked a bit later, config space > > can be assumed to be accessible. > > I don't quite follow this. In this path, pci_power_up() changes a > device from some low-power state to D0. If the device was in D3hot or > D3cold, we must delay at least 10ms before any access to it (PCIe > r7.0, sec 5.9). D3hot and D3cold are different in this respect, so using "or" here is inaccurate and confusing. Accessing the config space of a device in D3hot is entirely correct and doesn't require any delay. Accessing the config space of a device in D3cold is questionable because the device may not be accessible, but then the host bridge should just fail the access. The 10 ms delay is after attempting to program the device into D0 from D3hot (or the other way around) and it is observed as required. > pci_power_up() doesn't do any delay before the PMCSR read. It assumes that platform_pci_set_power_state() has run and either it has succeeded or the config space is not accessible which is when PCI_POSSIBLE_ERROR() will trigger. Unfortunately, there is no canonical way to verify that power has been restored to the device other than attempting to access it. > That part seems like a pre-existing issue even before this patch. I beg to differ. > If the PMCSR read returns PCI_POSSIBLE_ERROR(), pci_power_up() does > complain "Unable to change power state ... to D0" and return -EIO, but > pci_pm_power_up_and_verify_state() doesn't look at it, In fact, pci_update_current_state() changes the power state to D3cold if the config space is not accessible. > and pci_pm_default_resume_early() continues on to pci_restore_state(), so > it looks to me like we could try to restore state to an inaccessible > device. So the pci_restore_state() call could be avoided if the power state of the device was D3cold, but then the question is how much of a problem that is in practice. > We do call pci_dev_wait() in pci_pm_reset(), which does a D3hot -> D0 > transition; shouldn't we do the same in pci_power_up()? It would be good to make them consistent, but maybe the other way around?