tape-hardware compression stats

Chapman Flack <[email protected]>
Newsgroups gmane.comp.archivers.amanda.devel
Message-ID <[email protected]>
Hello -hackers,

If I use server compression and raw tape, it seems the curinfo
database gets populated with useful stats about different DLEs,
how well they compress, etc. Features like flush-threshold-dumped
and taperalgo largestfit ought to be fairly reliable, as they
know the compressed sizes of the filesets, and the real capacity
of the tape.

If I use only tape hardware compression, no statistics about
compressibility or final on-tape size get generated (I've just
inherited an installation that's running 3.2.3 and not the
latest version, so if my info is out of date and later versions
are different, I'd be happy to hear it!)

That makes the above-mentioned features sort of hit-or-miss,
because they are looking at the uncompressed holding-disk sizes,
and comparing to the manufacturer's hand-wavey "media capacity
with compression" figures, so Amanda might hold a bunch of
data until it's expected to be worth a tape, and then it
turns out to be half a tape, or two tapes, when actually
written to the media.

It might be that enough modern tape drives now support the
SCSI log sense compression pages (32h or 1Bh) ... and enough
popular OSes have something like the linux sg or solaris uscsi
generic driver features ... that it could be worthwhile for the
driver to be able to query after writing each fileset and
determine the actual on-tape size and compression ratio.

That's the part I think I could figure out on my own. What I'm
more curious about is, how much hackage would be needed in amanda
for stats like that, gathered by the tape device, to bubble back
up and be recorded somewhere useful, like with the curinfo stats?
Is that possible with existing protocols, or would a protocol
revision be needed first?

If the stats are gathered from the tape device, maybe they should
just go in curinfo exactly where the compressed sizes/ratios would
already go for server compression (since it is not likely that
both are in use).  Or would it be better to record them separately?

Am I right in assuming that if those stats were available, they
would get used by the existing logic that evaluates
flush-threshold-... and largestfit and so on, to make better
decisions? Or would that require additional development too?

Regards,
Chapman Flack
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.