Re: [PATCH v7 03/12] PCI: liveupdate: Track incoming preserved PCI devices
David Matlack <[email protected]> Tue, 21 Jul 2026 20:25:18 +0000
| Newsgroups | org.infradead.lists.kexec,org.kernel.vger.linux-doc,org.kernel.vger.linux-kernel,org.kernel.vger.linux-pci,org.kvack.linux-mm |
|---|---|
| Message-ID | <[email protected]> |
On 2026-07-21 05:55 PM, Pasha Tatashin wrote: > On 2026-07-20 16:07:47-07:00, David Matlack wrote: > > On Mon, Jul 20, 2026 at 3:44 PM Pasha Tatashin > > <[email protected]> wrote: > > > > > On 2026-07-20 14:54:51-07:00, David Matlack wrote: > > > > > > Thanks for the explanation. As I understand, the most straightforward > > > way to avoid holding the permanent reference is indeed to delete > > > dev->liveupdate.incoming entirely and perform an xarray lookup on every > > > access, like this: > > > > > > bool pci_liveupdate_is_incoming(struct pci_dev *dev) > > > { > > > ... > > > incoming = pci_liveupdate_flb_get_incoming(); > > > ... > > > dev_ser = xa_load(&incoming->xa, key); > > > ... > > > pci_liveupdate_flb_put_incoming(); > > > return dev_ser && dev_ser->refcount > 0; > > > } > > > > > > However, as you note, this is inefficient because it affects every > > > single device and adds lookup overhead to every access (not sure about > > > the actual cost though, xarray access is pretty fast!). > > > > > > We can, however, still avoid tinkering with the lifecycle of the FLB, > > > and instead treat dev->liveupdate.incoming as a "hint" that we validate > > > on access with a fast, liveness check: > > > > Can you tell me more about your concern about FLB lifetime? > > > > The lifetime of the FLB will not be affected by this reference unless > > there is a bug in the driver where it fails to call > > pci_liveupdate_finish() during it's file handler finish callback. > > From a design perspective, liveupdate_flb_get/put_incoming() is a > logical read-lock/unlock pair on the FLB data. We use a refcount for > optimization and sharing, but holding a get over a long asynchronous gap > (from boot-time device setup to driver probe) is essentially holding an > unbound lock. > > Unbound locks make it difficult to trace refcount leaks or debug > lifecycle issues. It's very standard for refcounts to be held for a long time, e.g. file refcounts and device refcounts. LUO FLBs themselves already have long-lived refcounts taken by each file that depends on themm, which aren't released until that file handler's finish callback runs. The PCI core is just doing the same thing. > > > > > > 1. At Setup: In pci_liveupdate_setup_device(), we do the xarray lookup > > > once, cache the pointer in dev->liveupdate.incoming, and immediately > > > call pci_liveupdate_flb_put_incoming(). We do not hold a permanent > > > reference. > > > > > > 2. On Access: When an accessor runs, instead of doing a full xarray > > > lookup, it just validates the cached pointer's liveness by temporarily > > > securing the FLB: > > > > > > static struct pci_flb_incoming *pci_liveupdate_get_incoming(struct pci_dev *dev) > > > { > > > struct pci_flb_incoming *incoming; > > > > > > incoming = pci_liveupdate_flb_get_incoming(); > > > if (!incoming) > > > return NULL; > > > > > > if (dev->liveupdate.incoming) > > > return incoming; > > > > > > pci_liveupdate_flb_put_incoming(); > > > return NULL; > > > } > > > > > > * If get_incoming() returns NULL (the FLB has already finished/freed), > > > the hint is invalid and the device is no longer incoming. > > > > This avoids the xarray lookup but still requires taking the incoming > > FLB mutex twice (once for get and once for put) on every access. And > > if there's no incoming PCI FLB, the LUO will iterate over all incoming > > FLBs under the mutex to find it. > > Can we do a fast-path check first? > > static struct pci_flb_incoming *pci_liveupdate_get_incoming(struct pci_dev *dev) > { > struct pci_flb_incoming *incoming; > > /* Fast-path to avoid unnecessary FLB querying */ > if (!dev->liveupdate.incoming) > return NULL; > > incoming = pci_liveupdate_flb_get_incoming(); > if (!incoming) > return NULL; > > /* Check again, now that FLB is acquired */ > if (dev->liveupdate.incoming) > return incoming; > > pci_liveupdate_flb_put_incoming(); > return NULL; > } > > This seems to gives us the best of both worlds: robust refcount hygiene > and a sane fast path. > > What do you think? It still seems like an overall worse approach to me. * From a performance perspective, the PCI core would still have to acquire the FLB mutex twice every time it needs to use the device's serialized state (once to acquire a new reference and once to release it). * From a locking perspective, the refcount doesn't actually protect against the device itself going through finish. So we still need the pci_liveupdate.rwsem. * From a hygiene perspective, each reference is held for less time yes, but there will more places in the code that need to acquire and release a refcount. That is more room for bugs to leak a refcount. So I'm not sure taking more small refcounts is an improvement. With the current approach there is just one refcount with a very clear lifetime (that aligns with the device file's refcount) and no extra LUO mutexes required to use the device's serialized state during device enumeration & setup.