[PATCH net v2 1/1] xsk: account ring allocations to RLIMIT_MEMLOCK
Zihan Xi <[email protected]>
| Newsgroups | org.kernel.vger.bpf,org.kernel.vger.netdev |
|---|---|
| Message-ID | <[email protected]> |
AF_XDP lets user space create RX, TX, fill and completion rings via
setsockopt() and mmap them into the process. xskq_create() allocates the
ring backing memory, but unlike XDP UMEM registration it never accounts
those pages against RLIMIT_MEMLOCK / user->locked_vm.
A process with CAP_NET_RAW in a user and network namespace can therefore
request very large rings and pin a large amount of kernel memory before
bind or any packet I/O. The later vmalloc_user() conversion changed the
allocation mechanism, but it did not create this missing resource boundary:
the unaccounted mmapable ring allocation already existed in the original
shared xskq_create() helper.
Fix this by accounting AF_XDP ring allocations with the existing
mm_account_pinned_pages() helper and unaccounting them when the queue is
destroyed. This puts ring memory under the same RLIMIT_MEMLOCK /
user->locked_vm model that AF_XDP already uses for UMEM pages, without
introducing AF_XDP-specific limits or socket lifetime coupling.
Fixes: 423f38329d26 ("xsk: add umem fill queue support and mmap")
Cc: [email protected]
Reported-by: Vega <[email protected]>
Assisted-by: Codex:gpt-5.4
Signed-off-by: Zihan Xi <[email protected]>
---
changes in v2:
- replace the socket optmem limit proposal with RLIMIT_MEMLOCK /
user->locked_vm accounting via mm_account_pinned_pages()
- drop the earlier socket/pool lifetime coupling changes and keep the
final code diff limited to xsk_queue.c and xsk_queue.h
- retarget Fixes to 423f38329d26, the original mmapable queue
allocation that introduced the missing resource boundary
- v1 Link: https://lore.kernel.org/all/[email protected]/
net/xdp/xsk_queue.c | 8 ++++++++
net/xdp/xsk_queue.h | 4 +++-
2 files changed, 11 insertions(+), 1 deletion(-)
diff --git a/net/xdp/xsk_queue.c b/net/xdp/xsk_queue.c
index 4dd01b7d858e..8be251096277 100644
--- a/net/xdp/xsk_queue.c
+++ b/net/xdp/xsk_queue.c
@@ -25,6 +25,7 @@ struct xsk_queue *xskq_create(u32 nentries, bool umem_queue)
{
struct xsk_queue *q;
size_t size;
+ int err;
q = kzalloc_obj(*q);
if (!q)
@@ -45,9 +46,15 @@ struct xsk_queue *xskq_create(u32 nentries, bool umem_queue)
}
size = PAGE_ALIGN(size);
+ err = mm_account_pinned_pages(&q->mmp, size);
+ if (err) {
+ kfree(q);
+ return NULL;
+ }
q->ring = vmalloc_user(size);
if (!q->ring) {
+ mm_unaccount_pinned_pages(&q->mmp);
kfree(q);
return NULL;
}
@@ -62,5 +69,6 @@ void xskq_destroy(struct xsk_queue *q)
return;
vfree(q->ring);
+ mm_unaccount_pinned_pages(&q->mmp);
kfree(q);
}
diff --git a/net/xdp/xsk_queue.h b/net/xdp/xsk_queue.h
index 1bc42c8902f4..b2e96c65f955 100644
--- a/net/xdp/xsk_queue.h
+++ b/net/xdp/xsk_queue.h
@@ -7,6 +7,7 @@
#define _LINUX_XSK_QUEUE_H
#include <linux/types.h>
+#include <linux/skbuff.h>
#include <linux/if_xdp.h>
#include <net/xdp_sock.h>
#include <net/xsk_buff_pool.h>
@@ -46,6 +47,7 @@ struct xsk_queue {
u64 invalid_descs;
u64 queue_empty_descs;
size_t ring_vmalloc_size;
+ struct mmpin mmp;
/* Mutual exclusion of the completion ring in the SKB mode.
* Protect: when sockets share a single cq when the same netdev
* and queue id is shared.
@@ -532,6 +534,6 @@ static inline u64 xskq_nb_queue_empty_descs(struct xsk_queue *q)
}
struct xsk_queue *xskq_create(u32 nentries, bool umem_queue);
-void xskq_destroy(struct xsk_queue *q_ops);
+void xskq_destroy(struct xsk_queue *q);
#endif /* _LINUX_XSK_QUEUE_H */
--
2.43.0