Re: [PATCH] hwmon: (corsair-psu) serialize debugfs access against hwmon
Wilken Gottwalt <[email protected]> Sun, 02 Aug 2026 12:57:46 +0000
| Newsgroups | org.kernel.vger.linux-hwmon,org.kernel.vger.linux-kernel |
|---|---|
| Message-ID | <[email protected]> |
On Sun, 2 Aug 2026 12:36:53 +0000 Ali Ahmet Memis <[email protected]> wrote: > corsairpsu_request() sends a rail select command and then the actual > read as two separate transfers, both going through the single shared > cmd_buffer and wait_completion in corsairpsu_usb_cmd(). The hwmon core > serializes its own callers, but the debugfs files call > corsairpsu_get_value() directly and never take that lock, so a debugfs > read can land between another reader's rail select and its value read. > > The result is a value from the wrong rail reported as the right one, > because corsairpsu_usb_cmd() only checks the command echo and both > transfers echo the command it expects. It can also make a caller consume > the reply meant for the other one, since raw_event() writes into the > shared buffer and completes whoever happens to be waiting. > > Locking was dropped in commit 4207069edbf0 ("hwmon: (corsair-psu) Rely > on subsystem locking") on the grounds that the subsystem serializes for > us, which holds for sysfs but not for these files. Take > the same lock in the debugfs paths that issue commands, using the guard > added in commit d1e720c7328e ("hwmon: Support guard() and scoped_guard > for subsystem locks"), as suggested in [1]. > > The lock cannot go into corsairpsu_request() itself: the hwmon core > already holds it across ->read, so every sysfs read would deadlock. > vendor_show() and product_show() only print strings cached during probe > and issue no command, and corsairpsu_get_criticals() and > corsairpsu_check_cmd_support() run before either interface is > registered, so none of them need it. > > [1] https://lore.kernel.org/all/[email protected]/ > > Fixes: 4207069edbf0 ("hwmon: (corsair-psu) Rely on subsystem locking") > Signed-off-by: Ali Ahmet Memis <[email protected]> > --- > This is the fix Guenter asked for in the May thread, written the way he > suggested there. Wilken's patch used a driver private mutex around > corsairpsu_request(); that thread stalled and the race is still present. > > Wilken, does this cover the chained command case you were worried about? > As far as I can tell it does: the whole select-rail plus read sequence > now runs under the same lock the hwmon core takes around ->read, so a > debugfs reader cannot land in the middle of one. If you had a case in > mind that this misses, I would rather hear it than guess. Yes, I think that is what Guenter asked me to test. There is actually a way to get all values from the PSU at once. You can chain together all the commands and everything supported should even fit into a single USB HID frame (64bytes). That would make everything a bit easier. Though, sorry that I did not go on with that. About a day after this someone put basically all my open source projects through an AI agent and since then I get bombarded with AI slop. I currently have not much energy (and fun) left doing my projects. I think I will test it in the next days. greetings Wilken > I have no Corsair PSU, so this is reasoned from the code rather than > measured on hardware. What I did check: > > - hwmon_lock() takes hwdev->lock, and the hwmon core takes the same > mutex around ->read and ->write, so this really does serialize the > two entry points > - the lock therefore cannot go into corsairpsu_request(), the sysfs > path would deadlock on itself > - probe registers the hwmon device before creating the debugfs files > and remove tears them down in the opposite order, so priv->hwmon_dev > is always valid inside a debugfs read > > Prior discussion: > https://lore.kernel.org/all/[email protected]/ > > drivers/hwmon/corsair-psu.c | 4 ++++ > 1 file changed, 4 insertions(+) > > diff --git a/drivers/hwmon/corsair-psu.c b/drivers/hwmon/corsair-psu.c > index ce958cdaef58..24100519cd83 100644 > --- a/drivers/hwmon/corsair-psu.c > +++ b/drivers/hwmon/corsair-psu.c > @@ -664,6 +664,8 @@ static void print_uptime(struct seq_file *seqf, u8 cmd) > long val; > int ret; > > + guard(hwmon_lock)(priv->hwmon_dev); > + > ret = corsairpsu_get_value(priv, cmd, 0, &val); > if (ret < 0) { > seq_puts(seqf, "N/A\n"); > @@ -730,6 +732,8 @@ static int ocpmode_show(struct seq_file *seqf, void *unused) > * getting of the value itself can also fail during this. Because of this every other > value > * than OCP_MULTI_RAIL can be considered as "single rail". > */ > + guard(hwmon_lock)(priv->hwmon_dev); > + > ret = corsairpsu_get_value(priv, PSU_CMD_OCPMODE, 0, &val); > if (ret < 0) > seq_puts(seqf, "N/A\n"); > > base-commit: 2d2338c93da79b3bfe4b6099a931d9468d539952