Re: [PATCH bpf-next v5 2/5] bpf: Add ksock kfuncs

Mahe Tardy <[email protected]>
Newsgroups org.kernel.vger.netdev,org.kernel.vger.bpf
Message-ID <[email protected]>
On Mon, Aug 10, 2026 at 09:24:38PM -0700, Kuniyuki Iwashima wrote:
> On Fri, Aug 7, 2026 at 10:15 AM Mahe Tardy <[email protected]> wrote:
> >
> > Add BPF kfuncs that allow BPF LSM programs to create and use sockets for
> > sending data. This provides a mechanism for BPF programs to emit
> > telemetry. For this first patch set, it's restricted to SOCK_DGRAM
> > socket types with IPPROTO_UDP protocol but could be easily extended to
> > SOCK_STREAM and IPPROTO_TCP in the future.
> >
> > The API consists of five kfuncs:
> >
> >   bpf_ksock_create()   - Create a socket (sleepable)
> >   bpf_ksock_connect()  - Connect socket to remote address (sleepable)
> >   bpf_ksock_send()     - Send data through the socket (sleepable)
> >   bpf_ksock_acquire()  - Acquire a reference to a socket context
> >   bpf_ksock_release()  - Release a reference (cleanup via
> >                          queue_rcu_work since sock_release sleeps)
> >

[...]

> > +
> > +/**
> > + * struct bpf_ksock - refcounted BPF kernel socket context
> > + * @sock:      The underlying kernel socket.
> > + * @usage:     Reference counter.
> > + * @rwork:     RCU work for deferred cleanup (sock_release may sleep).
> > + */
> > +struct bpf_ksock {
> > +       struct socket *sock;
> > +       refcount_t usage;
> > +       struct rcu_work rwork;
> > +};
> > +
> > +static void ksock_release_work_fn(struct work_struct *work)
> > +{
> > +       struct bpf_ksock *ks =
> > +               container_of(to_rcu_work(work), struct bpf_ksock, rwork);
> 
> nit: if you need respin:
> 
> struct bpf_ksock *ks;
> 
> ks = container_of(to_rcu_work(work), struct bpf_ksock, rwork);

sure!

> 
> > +
> > +       sock_release(ks->sock);
> > +       kfree(ks);
> > +}
> > +

[...]

> > +__bpf_kfunc struct bpf_ksock *
> > +bpf_ksock_create(const struct bpf_ksock_create_opts *opts, u32 opts__sz,
> > +                int *err__uninit)
> > +{
> > +       struct bpf_ksock_create_opts opts_copy;
> > +       struct bpf_ksock *ks;
> > +       int err;
> > +
> > +       /*
> > +        * sock_create() derives the network namespace, credentials, and cgroup
> > +        * from current. Kernel threads, including BPF workqueue callbacks, do
> > +        * not carry the context of the task that invoked the BPF program.
> > +        */
> > +       if (!bpf_ksock_has_user_task_context()) {
> > +               err = -EOPNOTSUPP;
> > +               goto err_out;
> > +       }
> > +
> > +       if (!opts || opts__sz != sizeof(struct bpf_ksock_create_opts)) {
> > +               err = -EINVAL;
> > +               goto err_out;
> > +       }
> > +
> > +       opts_copy = (struct bpf_ksock_create_opts){
> > +               .family = READ_ONCE(opts->family),
> 
> I'm not sure if we really want to care about such an arch, but
> since you mentioned unaligned access in bpf_ksock_connect(),
> is READ_OCNE() safe here ?

To my understanding, there's no issue regarding alignement here since
these four fields are __u8. The READ_ONCE is only used so that the
compiler does not do smart things like moving load or reloading a field
between time of check and time of use.

> 
> 
> > +               .type = READ_ONCE(opts->type),
> > +               .protocol = READ_ONCE(opts->protocol),
> > +               .reserved = READ_ONCE(opts->reserved),
> > +       };
> > +

[...]
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.