Re: [PATCH v4] svcrdma: Use contiguous pages for RDMA Read sink buffers
Mike Snitzer <[email protected]>
| Newsgroups | gmane.linux.drivers.rdma,gmane.linux.nfs |
|---|---|
| Message-ID | <[email protected]> |
On Thu, Mar 19, 2026 at 09:36:10AM -0400, Chuck Lever wrote: > From: Chuck Lever <[email protected]> > > svc_rdma_build_read_segment() constructs RDMA Read sink > buffers by consuming pages one-at-a-time from rq_pages[] > and building one bvec per page. A 64KB NFS READ payload > produces 16 separate bvecs, 16 DMA mappings, and > potentially multiple RDMA Read WRs (on platforms with > 4KB pages). > > A single higher-order allocation followed by split_page() > yields physically contiguous memory while preserving > per-page refcounts. A single bvec spanning the contiguous > range causes rdma_rw_ctx_init_bvec() to take the > rdma_rw_init_single_wr_bvec() fast path: one DMA mapping, > one SGE, one WR. > > The split sub-pages replace the original rq_pages[] entries, > so all downstream page tracking, completion handling, and > xdr_buf assembly remain unchanged. > > Allocation uses __GFP_NORETRY | __GFP_NOWARN and falls back > through decreasing orders. If even order-1 fails, the > existing per-page path handles the segment. > > When nr_pages is not a power of two, get_order() rounds up > and the allocation yields more pages than needed. The extra > split pages replace existing rq_pages[] entries (freed via > put_page() first), so there is no net increase in per- > request page consumption. Successive segments reuse the > same padding slots, preventing accumulation. The > rq_maxpages guard rejects any allocation that would > overrun the array, falling back to the per-page path. > Under memory pressure, __GFP_NORETRY causes the higher- > order allocation to fail without stalling. > > The contiguous path is attempted when the segment starts > page-aligned (rc_pageoff == 0) and spans at least two > pages. NFS WRITE segments carry application-modified byte > ranges of arbitrary length, so the optimization is not > restricted to power-of-two page counts. > > Signed-off-by: Chuck Lever <[email protected]> > --- > Changes since v3: > - Drop 1/3 - 3/3, they have already been reviewed and queued > - Incorporate hch's review comments > - Remove the #ifdef SZ_64K -- the logic works on those systems too > > > net/sunrpc/xprtrdma/svc_rdma_rw.c | 213 ++++++++++++++++++++++++++++++ > 1 file changed, 213 insertions(+) This patch, which landed during the 7.1 merge window as commit 18755b8c2f241 ("svcrdma: Use contiguous pages for RDMA Read sink buffers"), severely hurts RDMA performance when testing on very fast RDMA networking on x86_64. With this commit WRITE performance is "only" 24.2GB/s. Without this commit WRITE performance is 60.6GB/s. The cpu burn due to spinlock dominates the flamegraph that was collected. Chuck, I'll send you the flamegraph off-list (and can send it to anyone else who might be interested). Jon Flynn did the testing and we can request more info from him. We may have a window _now_ on the current Hammerspace testbed to try an incremental fix, but as of now simply reverting commit 18755b8c2f241 is as far as we got. Thanks, Mike