Re: [PATCH bpf-next v5 2/5] bpf: Add ksock kfuncs
Mahe Tardy <[email protected]>
| Newsgroups | org.kernel.vger.netdev,org.kernel.vger.bpf |
|---|---|
| Message-ID | <[email protected]> |
On Mon, Aug 10, 2026 at 09:24:38PM -0700, Kuniyuki Iwashima wrote: > On Fri, Aug 7, 2026 at 10:15 AM Mahe Tardy <[email protected]> wrote: > > > > Add BPF kfuncs that allow BPF LSM programs to create and use sockets for > > sending data. This provides a mechanism for BPF programs to emit > > telemetry. For this first patch set, it's restricted to SOCK_DGRAM > > socket types with IPPROTO_UDP protocol but could be easily extended to > > SOCK_STREAM and IPPROTO_TCP in the future. > > > > The API consists of five kfuncs: > > > > bpf_ksock_create() - Create a socket (sleepable) > > bpf_ksock_connect() - Connect socket to remote address (sleepable) > > bpf_ksock_send() - Send data through the socket (sleepable) > > bpf_ksock_acquire() - Acquire a reference to a socket context > > bpf_ksock_release() - Release a reference (cleanup via > > queue_rcu_work since sock_release sleeps) > > [...] > > + > > +/** > > + * struct bpf_ksock - refcounted BPF kernel socket context > > + * @sock: The underlying kernel socket. > > + * @usage: Reference counter. > > + * @rwork: RCU work for deferred cleanup (sock_release may sleep). > > + */ > > +struct bpf_ksock { > > + struct socket *sock; > > + refcount_t usage; > > + struct rcu_work rwork; > > +}; > > + > > +static void ksock_release_work_fn(struct work_struct *work) > > +{ > > + struct bpf_ksock *ks = > > + container_of(to_rcu_work(work), struct bpf_ksock, rwork); > > nit: if you need respin: > > struct bpf_ksock *ks; > > ks = container_of(to_rcu_work(work), struct bpf_ksock, rwork); sure! > > > + > > + sock_release(ks->sock); > > + kfree(ks); > > +} > > + [...] > > +__bpf_kfunc struct bpf_ksock * > > +bpf_ksock_create(const struct bpf_ksock_create_opts *opts, u32 opts__sz, > > + int *err__uninit) > > +{ > > + struct bpf_ksock_create_opts opts_copy; > > + struct bpf_ksock *ks; > > + int err; > > + > > + /* > > + * sock_create() derives the network namespace, credentials, and cgroup > > + * from current. Kernel threads, including BPF workqueue callbacks, do > > + * not carry the context of the task that invoked the BPF program. > > + */ > > + if (!bpf_ksock_has_user_task_context()) { > > + err = -EOPNOTSUPP; > > + goto err_out; > > + } > > + > > + if (!opts || opts__sz != sizeof(struct bpf_ksock_create_opts)) { > > + err = -EINVAL; > > + goto err_out; > > + } > > + > > + opts_copy = (struct bpf_ksock_create_opts){ > > + .family = READ_ONCE(opts->family), > > I'm not sure if we really want to care about such an arch, but > since you mentioned unaligned access in bpf_ksock_connect(), > is READ_OCNE() safe here ? To my understanding, there's no issue regarding alignement here since these four fields are __u8. The READ_ONCE is only used so that the compiler does not do smart things like moving load or reloading a field between time of check and time of use. > > > > + .type = READ_ONCE(opts->type), > > + .protocol = READ_ONCE(opts->protocol), > > + .reserved = READ_ONCE(opts->reserved), > > + }; > > + [...]