Re: Swap partition priority
Joshua Armstrong <[email protected]>
| Newsgroups | gmane.os.freebsd.devel.hackers |
|---|---|
| Message-ID | <[email protected]> |
I can do an experiment. I currently swap to a GEOM mirror. I can create an unmirrored fstab and take one disk offline and run the machine off a single disk and see if the swap writes require less memory for block IO. -Joshua > On Aug 18, 2026, at 10:44 AM, Mark Millard <[email protected]> wrote: > > On 8/18/26 08:21, Joshua Armstrong wrote: >> Short answer, yes - interleaving is hardcoded and the swapon(2) syscall >> only accepts the device node name (or file name). Here’s the doc header >> comment for the allocation code in sys/vm/swap_pager.c: >> >> 915 /* >> >> 916 * SWP_PAGER_GETSWAPSPACE() - allocate raw swap space >> >> 917 * >> >> 918 * Allocate swap for up to the requested number of pages. The >> >> 919 * starting swap block number (a page index) is returned or >> >> 920 * SWAPBLK_NONE if the allocation failed. >> >> 921 * >> >> 922 * Also has the side effect of advising that somebody made >> a mistake >> >> 923 * when they configured swap and didn't configure enough. >> >> 924 * >> >> 925 * This routine may not sleep. >> >> 926 * >> >> 927 * We allocate in round-robin fashion from the configured >> devices. >> >> 928 */ >> >> >> However, the suggestion to use a concatenated GEOM should work great >> because that’s a single device according to the swap and the GEOM >> allocator is (should be) entirely transparent to the swap strategy: >> >> >> Relevant snippet from swapgeom_strategy: >> >> >> 3212 bp->b_caller1 = bio; >> >> 3213 bio->bio_caller1 = sp; >> >> 3214 bio->bio_caller2 = bp; >> >> 3215 bio->bio_cmd = bp->b_iocmd; >> >> 3216 bio->bio_offset = (bp->b_blkno - sp->sw_first) * PAGE_SIZE; >> >> 3217 bio->bio_length = bp->b_bcount; >> >> 3218 bio->bio_done = swapgeom_done; >> >> 3219 bio->bio_flags |= BIO_SWAP; >> >> 3220 if (!buf_mapped(bp)) { >> >> 3221 bio->bio_ma = bp->b_pages; >> >> 3222 bio->bio_data = unmapped_buf; >> >> 3223 bio->bio_ma_offset = (vm_offset_t)bp->b_offset & >> PAGE_MASK; >> >> 3224 bio->bio_ma_n = bp->b_npages; >> >> 3225 bio->bio_flags |= BIO_UNMAPPED; >> >> 3226 } else { >> >> 3227 bio->bio_data = bp->b_data; >> >> 3228 bio->bio_ma = NULL; >> >> 3229 } >> >> 3230 g_io_request(bio, cp); >> >> 3231 return; >> >> >> >>> On Aug 18, 2026, at 9:00 AM, Joshua Armstrong <[email protected]> >>> wrote: >>> >>> I’ll throw my hat into this ring. >>> >>> As soon as I’m finished with my meeting, I’ll check the swap >>> management / virtual memory code to see if interleaving is hardcoded >>> or not. I believe the current (and long-standing, according to my >>> ancient copy of the Unix System Administration Handbook) behavior is >>> to interleave swap writes between configured devices for performance. >>> >>> I agree, however, that having a priority system would make more sense >>> in today’s world. I sometimes end up creating a temporary zvol and >>> using that as emergency extra swap if I’m building world, but I would >>> assume this has a performance penalty (though perhaps not massive) >>> versus swap partitions on bare disks. >>> >>> -Joshua >>> >>>> On Aug 18, 2026, at 3:48 AM, Miroslav Lachman <[email protected]> wrote: >>>> >>>> On 18/08/2026 09:49, Stéphane Rochoy wrote: >>>>> Andrea Venturoli <[email protected]> writes: >>>>>> Is it possible to set swap partition priority? >>>>>> According to swapon(8), the answer is no, but I wonder if the man page >>>>>> might be outdated, if patches are available, etc... >>>>>> >>>>>> The reason is I have SSDs and HDDs in the same system: I normally use >>>>>> swap on SSDs (for speed obviously), but sometimes the amount is not >>>>>> enough. >>>>>> So I'm ponderind adding swap on HDDS too, but this would slow down >>>>>> everyrting unless it's really used as last resort. >>>>> Hello, >>>>> Long ago I came accross a setup were the swap device was a GEOM mirror >>>>> (see gmirror(8)). AFAIK it was working except for vmcores: the kernel >>>>> was always using the first device which required to help savecore(8) a >>>>> bit. >>>>> So maybe there's some GEOM class matching your needs. >>>> There is geom_concat - see man gconcat. It may be possible to use it >>>> for swap (never tried it). >>>> >>>> Best regards >>>> Miroslav Lachman >>>> >>> >> > > Konstantin Belousov once wrote on the freebsd-arm list: > > QUOTE > . . . > swapfile write requires the write request to come through the filesystem > write path, which might require the filesystem to allocate more memory > and read some data. E.g. it is known that any ZFS write request > allocates memory, and that write request on large UFS file might require > allocating and reading an indirect block buffer to find the block number > of the written block, if the indirect block was not yet read. > > As result, swapfile swapping is more prone to the trivial and > unavoidable deadlocks where the pagedaemon thread, which produces free > memory, needs more free memory to make a progress. Swap write on the > raw partition over simple partitioning scheme directly over HBA are > usually safe, while e.g. zfs over geli over umass is the worst construction. > END QOUTE > > I do not know if the suggested use of GEOM is an example of that type of > issue getting involved. > > -- > === > Mark Millard > marklmi at yahoo.com