[MODERATED] Re: Additional sampling fun

"Luck, Tony" <[email protected]> Mon, 2 Mar 2020 17:03:51 -0800
Newsgroups org.kernel.lore.historical-speck
Message-ID <[email protected]>
On Fri, Feb 28, 2020 at 10:53:25PM +0100, speck for Thomas Gleixner wrote:
> speck for "Luck, Tony" <[email protected]> writes:
> > On Fri, Feb 28, 2020 at 06:44:48PM +0100, speck for Thomas Gleixner wrote:
> >> Have several cores with a 10k+ interrupts per second and if you're
> >> unlucky they start to contend, then the every 64th interrupt will be
> >> measurable quite prominent.
> >> 
> >> But I agree with Greg, that we can tackle this on LKML without
> >> mentioning that particular issue.
> >
> > That code really shouldn't ever have been using RDSEED (which is documented
> > as NOT scaling across invocations on multiple cores).
> 
> The only thing what the SDM says is:
> 
>   Under heavy load, with multiple cores executing RDSEED in parallel, it
>   is possible for the demand of random numbers by software
>   processes/threads to exceed the rate at which the random number
>   generator hardware can supply them. This will lead to the RDSEED
>   instruction returning no data transitorily. The RDSEED instruction
>   indicates the occurrence of this situation by clearing the CF flag.
> 
> I does not tell that it's slow to return CF=0. And if it does the
> current code just ignores it and carries on.

The "IntelĀ® Digital Random Number Generator (DRNG) Software Implementation Guide"[1]
provides graphs in section 3.4.1 showing that RDRAND throughput scales
linearly with the number of threads up to some saturation level (at
about 10 threads in the graph ... but the footnote says that these
results are "estimated" and "for informational purposes only").

I had thought that there was an RDSEED graph, but that must have been
in some internal document.

> So the question is whether the original RDSEED is slow already in the
> contended case or if the ucode mitigation will make it so.

If you plot throughput for RDSEED you'll get a flat line. Whatever
byte rate out got from one thread is all there is.

-Tony

[1] https://software.intel.com/en-us/articles/intel-digital-random-number-generator-drng-software-implementation-guide