Re: [PATCH v5 16/17] s390/dasd: Re-enable discard support for ESE volumes
[email protected] Fri, 31 Jul 2026 16:54:14 +0000
| Newsgroups | org.kernel.vger.linux-s390 |
|---|---|
| Message-ID | <[email protected]> |
Thank you for your contribution! Sashiko AI review found 3 potential issue(s) to consider: - [Critical] Setting discard_granularity to a non-power-of-two value corrupts block layer discard math. - [High] Returning BLK_STS_NOTSUPP for a benign unaligned discard causes filesystems to permanently disable discard support. - [High] Potential divide-by-zero kernel panic in dasd_eckd_disc_limits if ext_bytes evaluates to 0. -- --- Patch [16]: [PATCH v5 16/17] s390/dasd: Re-enable discard support for ESE volumes --- commit 7e5698c663d8a9a7aaad634b522a9f242320381c Author: Stefan Haberland <[email protected]> s390/dasd: Re-enable discard support for ESE volumes Re-enable block-layer discard for ESE ECKD volumes, releasing thin space via release allocated space (RAS). This routes REQ_OP_DISCARD to a RAS release over the request's track range and defines limits using extent granularity. > diff --git a/drivers/s390/block/dasd.c b/drivers/s390/block/dasd.c > --- a/drivers/s390/block/dasd.c > +++ b/drivers/s390/block/dasd.c [ ... ] > @@ -3157,6 +3159,15 @@ static blk_status_t do_dasd_request(struct blk_mq_hw_ctx *hctx, > rc = BLK_STS_RESOURCE; > } else if (PTR_ERR(cqr) == -EINVAL) { > rc = BLK_STS_INVAL; > + } else if (PTR_ERR(cqr) == -EOPNOTSUPP) { > + /* > + * e.g. a discard that covers no whole extent. This is an > + * expected, benign outcome (fstrim ranges rarely align to > + * the large ESE extent granularity), so silence the > + * per-request block-layer error print for it. > + */ > + req->rq_flags |= RQF_QUIET; > + rc = BLK_STS_NOTSUPP; [Severity: High] Will this cause upper layers to completely disable discard support? When returning BLK_STS_NOTSUPP for a benign unaligned discard, filesystems like ext4 typically interpret this as the underlying device lacking discard support entirely. This causes them to clear their DISCARD mount flag, permanently disabling discard for the volume. Should this instead complete the request with BLK_STS_OK so that it acts as a silent no-op without breaking subsequent discard requests? > } else { > DBF_DEV_EVENT(DBF_ERR, basedev, > "CCW creation failed (rc=%ld) on request %p", [ ... ] > diff --git a/drivers/s390/block/dasd_eckd.c b/drivers/s390/block/dasd_eckd.c > --- a/drivers/s390/block/dasd_eckd.c > +++ b/drivers/s390/block/dasd_eckd.c [ ... ] > @@ -7708,6 +7767,49 @@ static unsigned int dasd_eckd_max_sectors(struct dasd_block *block) > return DASD_ECKD_MAX_BLOCKS << block->s2b_shift; > } > > +/* > + * Discard on ECKD releases space through RAS, which works on whole extents. > + * Advertise extent granularity so the block layer only sends extent-aligned > + * discards (avoiding partially specified extents), and only for volumes on ESE > + * hardware. Non-ESE devices are left without discard limits. > + */ > +static void dasd_eckd_disc_limits(struct dasd_block *block, > + struct queue_limits *lim) > +{ > + struct dasd_device *device = block->base; > + struct dasd_eckd_private *private = device->private; > + unsigned int logical_block_size = block->bp_block; > + unsigned int max_discard_sectors, max_bytes, ext_bytes; > + int recs_per_trk, trks_per_cyl, ext_limit, ext_size; > + > + if (!dasd_eckd_ese_capable(device) || dasd_eckd_ext_size(device) == 0) > + return; > + > + trks_per_cyl = private->rdc_data.trk_per_cyl; > + recs_per_trk = recs_per_track(&private->rdc_data, 0, logical_block_size); > + > + ext_size = dasd_eckd_ext_size(device); > + ext_limit = min(private->real_cyl / ext_size, DASD_ECKD_RAS_EXTS_MAX); > + ext_bytes = ext_size * trks_per_cyl * recs_per_trk * logical_block_size; > + max_bytes = UINT_MAX - (UINT_MAX % ext_bytes); [Severity: High] Can ext_bytes evaluate to zero here? If the hardware or hypervisor supplies malformed RDC data, or if recs_per_trk evaluates to 0 due to an unrecognized device type, ext_bytes will be zero. Would the subsequent modulo arithmetic (UINT_MAX % ext_bytes) then trigger a divide-by-zero panic during device initialization? > + if (max_bytes / ext_bytes > ext_limit) > + max_bytes = ext_bytes * ext_limit; > + > + max_discard_sectors = max_bytes / 512; > + > + lim->max_hw_discard_sectors = max_discard_sectors; > + /* > + * ext_bytes is the hardware extent size and is not a power of two, so > + * the block layer's power-of-two round_up()/round_down() alignment > + * helpers compute it only approximately. That is a hint, not a > + * correctness requirement: RAS safety is enforced in the CCW builder, > + * which rounds the range inward to whole extents and rejects a request > + * that covers no whole extent, so a misaligned range is never > + * over-released. At worst a few sub-extent discards are declined. > + */ > + lim->discard_granularity = ext_bytes; [Severity: Critical] Is it safe to use a non-power-of-two value for discard_granularity here? The block layer strictly requires discard_granularity to be a power of two because it relies on bitwise macros to compute discard alignments (such as sector &= ~(granularity >> SECTOR_SHIFT - 1) in __blkdev_issue_discard). Could supplying a non-power-of-two value break these assumptions and cause the block layer to compute incorrect sector offsets and lengths, potentially discarding live filesystem data instead of free space? > +} -- Sashiko AI review ยท https://sashiko.dev/#/patchset/[email protected]?part=16