[PATCH] block: skip redundant flush for O_DSYNC direct writes

Zhenxian Ma <[email protected]>
Newsgroups org.kernel.vger.linux-block
Message-ID <[email protected]>
For an O_DIRECT | O_DSYNC write to a block device, dio_bio_write_op()
adds REQ_FUA to the bio, so the data is durable when the direct-IO
returns.  blkdev_write_iter() then falls through to
generic_write_sync(), which for O_DSYNC turns into an additional
REQ_PREFLUSH.  That flush is redundant: the FUA write already made
the data persistent.

Skip generic_write_sync() when the direct-IO path already provided
durability via FUA.  The buffered and direct_write_fallback() paths
are unchanged.

Measured on a Seagate ST20000NM007D (20 TB, 7200 rpm, fua=1,
write_cache=write back), Linux v7.2.0-rc7, single-threaded pwrite()
loop opening the raw block device with O_WRONLY | O_DIRECT | O_DSYNC,
4 KiB writes for 60 s:

  Sequential 4 KiB writes:
                            baseline    patched
    IOPS                       119.7     7497.0
    avg latency (us)            8357        133
    p50 latency (us)            8346        127
    p99 latency (us)            8368        395
    p99.9 latency (us)          8728        569

  Random 4 KiB writes (100 GiB span):
                            baseline    patched
    IOPS                       156.1      666.4
    avg latency (us)            6405       1500
    p50 latency (us)            6186       1450
    p99 latency (us)           16133       2285
    p99.9 latency (us)         17250       9916

Signed-off-by: Zhenxian Ma <[email protected]>
Signed-off-by: Zhenxian Ma <[email protected]>
---
 block/fops.c | 9 +++++++--
 1 file changed, 7 insertions(+), 2 deletions(-)

diff --git a/block/fops.c b/block/fops.c
index 15783a6180de..9c34a372b476 100644
--- a/block/fops.c
+++ b/block/fops.c
@@ -726,6 +726,7 @@ static ssize_t blkdev_write_iter(struct kiocb *iocb, struct iov_iter *from)
 	bool atomic = iocb->ki_flags & IOCB_ATOMIC;
 	loff_t size = bdev_nr_bytes(bdev);
 	size_t shorted = 0;
+	bool dio_fua_done = false;
 	ssize_t ret;
 
 	if (bdev_read_only(bdev))
@@ -763,9 +764,13 @@ static ssize_t blkdev_write_iter(struct kiocb *iocb, struct iov_iter *from)
 
 	if (iocb->ki_flags & IOCB_DIRECT) {
 		ret = blkdev_direct_write(iocb, from);
-		if (ret >= 0 && iov_iter_count(from))
+		if (ret >= 0 && iov_iter_count(from)) {
 			ret = direct_write_fallback(iocb, from, ret,
 					blkdev_buffered_write(iocb, from));
+		} else if (ret > 0 && iocb_is_dsync(iocb)) {
+			/* FUA from dio_bio_write_op() already made it durable */
+			dio_fua_done = true;
+		}
 	} else {
 		/*
 		 * Take i_rwsem and invalidate_lock to avoid racing with
@@ -777,7 +782,7 @@ static ssize_t blkdev_write_iter(struct kiocb *iocb, struct iov_iter *from)
 		inode_unlock_shared(bd_inode);
 	}
 
-	if (ret > 0)
+	if (ret > 0 && !dio_fua_done)
 		ret = generic_write_sync(iocb, ret);
 	iov_iter_reexpand(from, iov_iter_count(from) + shorted);
 	return ret;
-- 
2.43.5
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.