Thank you for your contribution! Sashiko AI review found 3 potential issue(s) to consider:
Pre-existing issues:
- [Critical] Disabling XDP triggers an out-of-bounds array access due to the old BPF program remaining active while the TX arrays are shrunk.
- [High] Device resets permanently break AF_XDP sockets because the driver zeroes the `xsk_pools` bitmask.
- [High] `gve_xsk_pool_enable` causes a data race by modifying the XDP memory model concurrently with the NAPI polling context.
--
--- Patch [2]: [PATCH net-next v4 02/15] gve: refactor initialization with helper functions ---
commit f4310d54d17a7a878691c693aea2b3ce23d39802
Author: Harshitha Ramamurthy <[email protected]>
gve: refactor initialization with helper functions
In the interest of commonizing code, refactor gve_probe()
and gve_init_priv() with a few helper functions that can
be expanded and utilized in upcoming patches that add the
mailbox ABI to the driver.
> diff --git a/drivers/net/ethernet/google/gve/gve_main.c b/drivers/net/ethernet/google/gve/gve_main.c
> index 30bf6df4ebc58..37e6205e2eefd 100644
> --- a/drivers/net/ethernet/google/gve/gve_main.c
> +++ b/drivers/net/ethernet/google/gve/gve_main.c
> @@ -2400,7 +2400,6 @@ static const struct xdp_metadata_ops gve_xdp_metadata_ops = {
>
> static int gve_init_priv(struct gve_priv *priv, bool skip_describe_device)
> {
[Severity: High]
This is a pre-existing issue, but could device resets permanently break AF_XDP
sockets because the driver zeroes the xsk_pools bitmask?
When a device reset occurs (e.g., due to timeout, adminq failure, or MTU
change), it triggers gve_teardown_priv_resources() which frees the bitmask:
bitmap_free(priv->xsk_pools);
priv->xsk_pools = NULL;
During recovery, the initialization path reallocates a fresh zeroed
xsk_pools bitmap (e.g., via priv->xsk_pools = bitmap_zalloc(...)),
permanently losing the previous state. Because the driver never restores the
xsk_pools state from the core net_device, AF_XDP sockets silently stop
working after a reset.
[Severity: Critical]
This is a pre-existing issue, but is there a risk of an out-of-bounds
array access when disabling XDP?
In gve_set_xdp(), the xdp_prog pointer is updated after the queues have been
restarted with a shrunken TX array:
err = gve_configure_rings_xdp(priv, 0);
if (err)
goto out;
WRITE_ONCE(priv->xdp_prog, prog);
If a local privileged user disables XDP, gve_set_xdp() calls
gve_configure_rings_xdp(priv, 0), which reallocates priv->tx without XDP
queues and restarts NAPI polling. Since priv->xdp_prog is updated only after
NAPI restarts, the active NAPI poll can execute the old BPF program. If the
program returns XDP_TX, it accesses priv->tx[tx_qid] where tx_qid exceeds
the new, smaller array size, leading to memory corruption.
[Severity: High]
This is a pre-existing issue, but does gve_xsk_pool_enable() cause a data
race by modifying the XDP memory model concurrently with the NAPI polling
context?
In gve_xsk_pool_enable(), it registers an AF_XDP pool by modifying the XDP
memory model and rx->xsk_pool without first disabling NAPI polling:
err = xdp_rxq_info_reg_mem_model(&rx->xdp_rxq, MEM_TYPE_XSK_BUFF_POOL,
pool);
if (err) {
...
}
rx->xsk_pool = pool;
If NAPI receives a packet in the race window, it may initialize the XDP
buffer with MEM_TYPE_XSK_BUFF_POOL but allocate memory from the standard page
pool because rx->xsk_pool was observed as NULL. When freed, xsk_buff_free()
would be incorrectly called on a standard page, causing type confusion and
memory corruption.
--
Sashiko AI review · https://sashiko.dev/#/patchset/[email protected]?part=2
lmpx.com only provides a reader for public news (NNTP) servers. It is not
affiliated with the servers or forums shown here and is not responsible for
the content of articles, which is written by their respective authors.