Re: [RFC] bpf: account ring buffer backing pages separately from Lost RAM
"David Hildenbrand (Arm)" <[email protected]>
| Newsgroups | gmane.linux.file-systems,gmane.linux.kernel.bpf,gmane.linux.kernel.mm,gmane.linux.kernel |
|---|---|
| Message-ID | <[email protected]> |
On 8/15/26 11:18, Xiang Gao wrote: > Hi, Hi, > > I would like to discuss accounting BPF ring buffer backing pages in > system-wide memory reports. > > BPF ring buffers allocate their data and metadata as order-0 pages directly > from the buddy allocator, and then map those pages with vmap(). I assume there is a reason the slab isn't used, right? Are these pages mapped into user space such that page->mapcount would get used? Can you point me at relevant code? > > Because vmap() maps caller-owned pages, these backing pages are not counted > by VmallocUsed. They are also not slab pages. As a result, most BPF ring > buffer memory is not represented by an existing named /proc/meminfo category > and appears as Lost RAM in Android memory reports. > > We measured this on an Android 6.18 kernel. > > Test case: > > 32 BPF ring buffer maps > 16 MiB data area per map > 512 MiB total data area > > Observed changes: > > Lost RAM: approximately +529 MiB > VmallocUsed: approximately +2 MiB > Slab: approximately unchanged > > After destroying all maps, the values returned close to baseline. > > The question is whether the kernel should expose the unique physical backing > pages of live BPF ring buffers through a dedicated global counter and a > /proc/meminfo entry, for example: > > BpfRingbuf: <value in kB> This looks a bit too specific for my taste. And I think we should try to no inflate these statistics here too much. > > The proposed counter would include: > > * ring buffer data pages; > * metadata pages; > * consumer and producer position pages. > > It would exclude: > > * the second virtual mapping of data pages; > * the pages[] pointer array; > * map metadata allocations; > * vmap page tables. > > The goal is to account for the currently unclassified direct backing pages. > Slab- and vmalloc-backed auxiliary allocations are already represented by > existing memory categories and should not be counted again. > > A possible implementation is an NR_BPF_RINGBUF vmstat counter maintained by > the ring buffer allocation and free paths, with the aggregate exposed through > /proc/meminfo. > > Questions: > > 1. Is a dedicated BPF ring buffer counter appropriate? I don't think so. See [1] where we just had the same discussion for tracing buffers. For them, Steve [2] had an idea on how to expose them more fine-grained and tracing specific. [1] https://lore.kernel.org/r/[email protected] [2] https://lore.kernel.org/r/[email protected] > 2. Should this be represented as an NR_* vmstat counter? I don't think so. > 3. Is /proc/meminfo an acceptable interface for this information? Again, I don't think so. "Lost RAM" really is just "excessive memory allocated by some other subsystem". I agree that some users might want to figure out what is consuming that much memory, but I don't think growing /proc/meminfo in that way is really what we want. > 4. Is counting only unique physical backing pages the correct accounting unit? I'd assume the "It would exclude" part above should not be accounted there, if that's what you mean. -- Cheers, David