Re: [PATCH RFC 0/4] nfsd: per-client fair-queue dispatch
NeilBrown <[email protected]>
| Newsgroups | gmane.linux.nfs |
|---|---|
| Message-ID | <[email protected]> |
On Thu, 04 Jun 2026, Benjamin Coddington wrote: > knfsd dispatches ready transports from a single per-pool FIFO, so a > client's share of nfsd service scales with the number of connections it > holds rather than being shared per client. A client that opens many > connections (nconnect, or a farm of data movers) starves other clients > on the same server purely by out-numbering them in sockets. > > I measured this with a load generator that pins each request to a fixed > service time and does no filesystem work, so that nfsd thread-time is the > only scarce resource (8 threads, 10ms/op, ~648 ops/s pool ceiling). A > greedy client opens K connections alongside one single-connection > interactive client. > > NFSv3, dispatch as it is today: > > greedy K greedy share interactive ops/s > 1 50% 241 > 4 80% 129 > 8 89% 72 > 16 94% 38 > > The interactive client's share tracks 1/(K+1) and its throughput falls > roughly 6x while it does nothing different. NFSv4.1 behaves identically > (89% greedy at K=8) even when the greedy connections are bound to a > single session, because the dispatch decision is below the NFS version. > > The same NFSv4.1 workload with fair queueing enabled: > > greedy K greedy share interactive ops/s > 8 72% 182 > 16 73% 177 > 32 70% 193 > > The greedy client's share no longer climbs with its connection count and > the interactive client recovers (72 -> 182 ops/s at K=8). Aggregate > throughput is unchanged: the T/D pool ceiling is the same with fair > queueing on and off. The split does not reach 50/50 because a single > interactive connection is bounded by its request window and by XPT_BUSY > serialising one transport; with a deeper window it reaches ~59/41. > > The approach: > > - sunrpc grows an opaque per-transport fairness key (patch 1), with a > default derived from the source address (the source port is excluded > so a client's several connections share one key), and an opt-in > per-pool scheduler that buckets ready transports by that key and > dispatches round-robin across keys (patch 2). When it is disabled, > which is the default, the existing lockless FIFO path is unchanged. > > - nfsd gains a "fairq" module parameter to turn it on (patch 3) and > stamps the NFSv4.1 clientid as the key when a connection binds to a > session (patch 4), so all of a client's connections share one key. > NFSv3 uses the source-address default. > > This is an RFC; a few questions for the list: > > - Unit of fairness: clientid (used here) or session? Earlier > discussion leaned toward exploring per-session. > > - Mechanism: a fixed bucket hash under a per-pool spinlock taken only > on the opt-in path, versus a lockless or per-flow structure. I'm not keen on the hash bucket approach, though I can't clearly say why. I'm also not keen on the opt-in design. I don't like asking the admin to tune performance. We should always provide optimal performance. I imagine creating an object which represents a client - using IP address or possibly v4 client id. These may well be located using a hash table (rhashtable?), but that is peripheral to the design. Each xprt has an associated client which can be changed at any time (nfsd could change to a v4.1-client). Each client has a lwq of xprts and is part of an lwq of clients. To enqueue an xprt, we grab the client, enqueue to that, then if necessary enqueue the client. To dequeue next, we dequeue the first client, dequeue the first xprt, then optionally enqueue the client again. So this would be slightly more work than the current (2 dequeues instead of 1) but I think that might be acceptable. clients would be refcounted by xprts and probably rcu-freed. > > - Would a per-client in-flight cap be preferable to proportional fair > queueing? A per-client cap would only be ok if the admin didn't have to tune it. So the client would need to get some sort of feed-back from the server so that it knows when it is pushing too hard. If we were still using UDP we could possibly use packet-loss for that feed-back, but we aren't and don't want to. With v4.1 we can of course use the slot based flow control and I think we should if we can agree on a good design. With v3 I don't think there is any way to get the needed feed-back Thanks, NeilBrown > > The measurement used a debug-only filehandle-latency hook that is not > part of this series. > > Benjamin Coddington (4): > sunrpc: add a per-transport fairness key to svc_xprt > sunrpc: dispatch ready transports fairly per client > nfsd: add a fairq module parameter > nfsd: key NFSv4.1 connections by clientid for fair queueing > > fs/nfsd/nfs4state.c | 17 +++ > fs/nfsd/nfssvc.c | 19 +++ > include/linux/sunrpc/svc.h | 5 + > include/linux/sunrpc/svc_xprt.h | 46 ++++++- > net/sunrpc/svc.c | 2 + > net/sunrpc/svc_xprt.c | 216 +++++++++++++++++++++++++++++++- > 6 files changed, 302 insertions(+), 3 deletions(-) > > > base-commit: e7ca66ba17f1b5e4ecbb29b9c3c4a31aa062bed0 > -- > 2.53.0 > >