Re: [PATCH 1/2] crypto: pcrypt - Remove pcrypt
Hendrik Donner <[email protected]> Fri, 24 Jul 2026 20:19:02 +0200
| Newsgroups | org.kernel.vger.linux-crypto,org.kernel.vger.linux-kernel,org.kernel.vger.netdev |
|---|---|
| Message-ID | <[email protected]> |
Hallo, On 7/22/26 21:40, Eric Biggers wrote: > On Wed, Jul 22, 2026 at 06:12:16PM +0200, Hendrik Donner wrote: >> Hello, >> >> On 7/21/26 21:50, Eric Biggers wrote: >>> On Tue, Jul 21, 2026 at 08:59:21PM +0200, Hendrik Donner wrote: >>>> Hello, >>>> >>>> On 7/14/26 00:32, Eric Biggers wrote: >>>>> pcrypt was originally intended to improve IPsec performance. However, >>>>> it's no longer useful for that. Reports from the rare cases that anyone >>>>> has actually tried to use it over the years indicate that it actually >>>>> reduces IPsec performance, e.g.: >>>>> >>>>> * https://github.com/libreswan/libreswan/wiki/Internals:-Cryptographic-Acceleration#obsoleted-ipsec-accelerations >>>>> * https://users.strongswan.narkive.com/liqTaTq8/strongswan-problem-with-pcrypt >>>>> * https://unix.stackexchange.com/questions/594336/ipsec-multithreading-via-pcrypt-worse-than-single-thread >>>>> >>>>> It's also undocumented and quite difficult to actually use. Its design >>>>> is also broken, in that any unprivileged program can enable pcrypt >>>>> systemwide at any time (by instantiating it using AF_ALG). >>>>> >>>>> Meanwhile, pcrypt has been a regular source of bugs, including at least >>>>> four that have received CVEs. >>>>> >>>>> Let's just remove it. No one seems to care about it anymore other than >>>>> people looking for vulnerabilities. >>>>> >>>> >>>> my company is a user. We have a hardware platform based on an IMX6 SoC >>>> using IPSec and configure pcrypt using crconf. Current performance >>>> difference: >>>> >>>> iperf3 -c <IP> --time 60 -R >>>> >>>> pcrypt: >>>> Download: 107 Mbits/sec >>>> >>>> No pcrypt: >>>> Download: 59.3 Mbits/sec >>>> >>>> iperf3 -c <IP> --time 60 >>>> >>>> pcrypt: >>>> Upload: 65.9 Mbits/sec >>>> >>>> No pcrypt: >>>> Upload: 52.0 Mbits/sec >>>> >>>> The relevant crypto templates are configured in early userspace and >>>> since i got curious, that has been the case since 2017. >>>> >>>> Mostly using >>>> >>>> pcrypt(gcm_base(ctr-aes-neonbs,ghash-generic)) >>>> >>>> nowadays, AES-CBC in the past/as a fallback option. >>>> >>>> So at least on some platforms there is still a significant performance >>>> boots, at least for downloads in this case. >>> >>> Thanks for bringing up your use case. >>> >>> Have you looked into alternative solutions such as Receive Side Scaling >>> (https://docs.kernel.org/networking/scaling.html#rss-receive-side-scaling)? >>> AFAIK it's not just the crypto performance that makes pcrypt unnecessary >>> these days, but also the design of the networking layer. >> >> I'm looking into this more, the IMX.6 is a bit limited with IRQ handling and >> queue distribution. > > Thanks! Maybe Steffen and the other IPsec folks would have some advice > too. > so i'm now on 7.1.4 with PCI: imx6: Keep i.MX6 Root Port MSI/MSI-X Capabilities with iMSI-RX to work around hardware bug on top to be able to tune queue settings. And to have a working ethernet in the first place, without the patch the NETDEV WATCHDOG resets the card all the time due to queues stalling. But now more than 1 CPU are serving IRQs. With pcrypt (seqiv(rfc4106(pcrypt(gcm_base(ctr-aes-neonbs,ghash-lib))))): Upload: [ 4] 0.00-60.00 sec 901 MBytes 126 Mbits/sec Download: [ 4] 0.00-60.00 sec 1.23 GBytes 177 Mbits/sec Without pcrypt (seqiv(rfc4106(gcm_base(ctr-aes-neonbs,ghash-lib)))): Upload: [ 4] 0.00-60.00 sec 679 MBytes 94.9 Mbits/sec Download: [ 4] 0.00-60.00 sec 674 MBytes 94.3 Mbits/sec So counterintuitively pcrypt matters more again. I repeated the tests a few times, those numbers are fairly representative. Every run is over a 60 sec window. Regards, Hendrik >>> I understand that i.MX6 doesn't have the ARMv8 crypto extensions. >>> However, surely you could at least use the NEON-optimized GHASH code? >>> Is there a reason you're not using it? >> >> I think it was historically not working well for us, retested: >> >> NEON GHASH with pcrypt: >> >> Download: 116 Mbits/sec >> Upload: 74.3 Mbits/sec >> >> NEON GHASH baseline: >> >> Download: 90.7 Mbits/sec >> Upload: 60.8 Mbits/sec >> >> Looks better, still a ~15 Mbit improvement with pcrypt. >> >> I want to point out that without IPSec our baseline is: >> >> Download: 942 Mbits/sec >> Upload: 942 Mbits/sec >> >> Intel IGB ethernet. > > That already shows that about two-thirds of the improvement you were > getting from pcrypt can be gotten just by using the correct GHASH > implementation for the platform (59.3 => 90.7 download, vs 59.3 => 107; > and 52.0 => 60.8 upload, vs 52.0 => 65.9). > > And with that new baseline, for downloads, pcrypt adds just 28% more > throughput (90.7 => 116) rather than 80% as it did before (59.3 => 107). > > It seems clear that the usefulness of pcrypt rapidly decreases as the > actual crypto gets faster. > > Note: we've been enabling crypto optimizations by default in recent > kernels, so that people can no longer use the generic code by accident. > For example in v7.1 and later, GHASH optimizations are always enabled. > > I'm working on further optimizations to the AES-GCM code as well, > specifically implementing AES-GCM directly on all platforms without the > inefficient gcm_base template that is being used here. > > So while maybe pcrypt does still help a bit on this platform for now, > the approach does seem quite dated and largely a workaround for > inefficiencies elsewhere in the stack (including systems where the > optimized crypto code is accidentally not enabled). > > - Eric