[PATCH net-next v2 0/5] net/rds: own the fastpath locks across connection teardown

Allison Henderson <[email protected]>
Newsgroups org.kernel.vger.linux-rdma,org.kernel.vger.netdev
Message-ID <[email protected]>
Hi all,

This is v2 of the follow-up set to "net/rds: Bug fix ports, part 2"
[1] (v1 of this set is at [2]).  During review of part 2, the later
half of that series needed more work than a respin, so it was split
off into this set together with the companion fixes identified along
the way.

RDS connection teardown quiesces the transmit and receive-refill fast
paths by waiting for the RDS_IN_XMIT/RDS_RECV_REFILL bits to be
sampled clear.  Sampling a bit clear is not owning it: the fast path
can re-take its bit right after the wait returns and then run
concurrently with the transport shutdown and the send-state reset.
Oracle UEK closed this by making teardown acquire the bits as locks
("rds: Make sure transmit path and connection tear-down does not run
concurrently"); patches 4 and 5 do the same for the two
rds_send_path_reset() call sites upstream.  These two are effectively
v3 of patches 4 and 3 of "net/rds: Bug fix ports, part 2" [1].

Making teardown block on the bits as locks promotes three latent
ordering bugs from rare to load-bearing, so they are fixed first:

  Patch 1: release_in_xmit() checks waitqueue_active() after
  clear_bit_unlock(), which does not order that read; the wake-up of
  the (now uninterruptible, untimed) teardown wait can be lost.  Use
  wq_has_sleeper().

  Patch 2: rds_conn_path_reset() wipes the whole cp_flags word with a
  plain store.  Once teardown owns bits in that word across the
  reset, a blanket store would end lock ownership early - and it
  already races atomic RMWs on the same word today.  Clear the bits
  the reset is responsible for individually, as Oracle UEK also does.

  Patch 3: rds_tcp_reset_callbacks() stores RDS_CONN_RESETTING
  unconditionally, which can overwrite the RDS_CONN_ERROR or
  RDS_CONN_DISCONNECTING of a shutdown already in progress on the
  same path and send that shutdown through an extra drop cycle.  Once
  the accept path can park for the duration of a teardown (patch 5)
  that window widens, so make the transition conditional first, as
  Oracle UEK does.

With those in place, patch 4 converts rds_tcp_reset_callbacks() from
waiting on RDS_IN_XMIT to acquiring it, holding it across the socket
swap and rds_send_path_reset(), and patch 5 has rds_conn_shutdown()
hold both bit locks across the transport shutdown and path reset.
The order matters: with the accept path owning the lock first, no
intermediate commit leaves it resuming on a socket pointer that a
lock-holding teardown has already released.

[PATCH net-next 1/5] net/rds: use wq_has_sleeper() in release_in_xmit()
  Restore full barrier before wake-up checks in release_in_xmit()

[PATCH net-next 2/5] net/rds: clear cp_flags bits individually in rds_conn_path_reset()
  Partial port of commit d04896037223 ("net/rds: Preserve essential connection state flags")
  https://github.com/oracle/linux-uek/commit/d04896037223

[PATCH net-next 3/5] net/rds: tcp: don't force RDS_CONN_RESETTING over a concurrent shutdown
  Port of commit 72c176a1d9ac ("net/rds: Don't force state RDS_CONN_RESETTING")
  https://github.com/oracle/linux-uek/commit/72c176a1d9ac

[PATCH net-next 4/5] net/rds: acquire RDS_IN_XMIT in rds_tcp_reset_callbacks()
  Extend the port in patch 5 to the second rds_send_path_reset() call site

[PATCH net-next 5/5] net/rds: acquire the fastpath locks in rds_conn_shutdown()
  Port of commit 2b8aaa4f163b ("rds: Make sure transmit path and connection tear-down does not run concurrently")
  https://github.com/oracle/linux-uek/commit/2b8aaa4f163b

Changes since v1 [2]:
  - New patch 3, porting the UEK fix for the accept-vs-shutdown state
    race; the RESETTING store is now a conditional transition.
  - Patches 4 and 5 (formerly 4 and 3) reordered so that
    rds_tcp_reset_callbacks() owns RDS_IN_XMIT before
    rds_conn_shutdown() does; no intermediate commit is a bisect
    hazard.
  - Patch 2 gains a Fixes tag; patch 4 refreshes the stale block
    comment above rds_tcp_reset_callbacks(); patch 5's changelog
    describes the accept-side waiter parking on krdsd for the
    duration of a TCP teardown and why that is preferable to the race.

Questions and comments appreciated!

Thanks,
Allison

[1] https://lore.kernel.org/netdev/[email protected]/
[2] https://lore.kernel.org/netdev/[email protected]/


Allison Henderson (3):
  net/rds: use wq_has_sleeper() in release_in_xmit()
  net/rds: clear cp_flags bits individually in rds_conn_path_reset()
  net/rds: acquire RDS_IN_XMIT in rds_tcp_reset_callbacks()

Gerd Rausch (1):
  net/rds: tcp: don't force RDS_CONN_RESETTING over a concurrent
    shutdown

Håkon Bugge (1):
  net/rds: acquire the fastpath locks in rds_conn_shutdown()

 net/rds/connection.c | 32 +++++++++++++++++--
 net/rds/send.c       | 12 +++++--
 net/rds/tcp.c        | 76 +++++++++++++++++++++++++++++---------------
 3 files changed, 88 insertions(+), 32 deletions(-)


base-commit: 3da8c3c8b8fa99505624b65ef590482f48e766b6
-- 
2.25.1
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.