Re: Create a catalog from the rpm database

Denis Corbin <[email protected]> Mon, 27 Jan 2014 21:18:29 +0100
Newsgroups gmane.comp.sysutils.backup.dar.general
Message-ID <[email protected]>
-----BEGIN PGP SIGNED MESSAGE-----
Hash: SHA1

On 27/01/2014 09:19, gulikoza wrote:
> Hello,

Hello,

> 
> I use dar extensively for backups, usually separating system (the
> linux itself) and content backups.

:)

> I store system backups also in my archive in order to have some
> history of the changed configurations and such.

Do you mean that the system backups take place inside the content
backup? Not sure to understand (... sorry its late here :@) )

to have the history of changes you might be using dar_manager, unless
I don't exactly what you need to do...

> 
> The problem is that this is becoming a big organizational
> nightmare. A simple update might change a lot of files, bloating
> backup size

Yes, I've noted that upgrading a package (.deb or .rpm) often
reinstall 90% of the file concerned by that package ... which dar
detects as changed and save them again. You can eventually ignore some
attributes modifications, like permission or ownership while the file
size is the same [--comparison-field option] but at the risk you miss
some updated files...

> and if you multiply that with a few dozen servers, checking each
> and every backup is quite a task.

Another possibility would be to upgrade a reference server with
packaging tools and take differential backup of it with dar that you
push to other servers, no?

> However, most of the stuff is rarely needed as files not 
> specifically changed are better tracked by rpm itself (having a
> list of installed packages + changes would allow you to rebuild the
> same system). I have thought of creating a script that would list
> rpm changed and untracked files and include only those in the
> system backup, but there are some exceptions. Some files might be
> changed on the disk, but not listed in rpm verify for instance the
> rpm database itself:
> 
> %verify(not md5 size mtime) %ghost %config(missingok,noreplace) 
> /var/lib/rpm/*
> 
> The files in /var/lib/rpm will not be shown as modified by rpm -V,
> but will be included in rpm -ql. Any script working only with the
> list of changed and untracked files, would fail to include these in
> the backup.
> 
> At first I thought of somehow interfacing dar directly with the rpm
> database but then I decided for a simpler approach. The attached
> script tries to create an empty root tree (sparse files) from the
> rpm database that can be used to create a dar catalogue (-A +, a
> snapshot backup) which in turn can be used as a reference for a
> system backup. Such a backup would be pretty lightweight and allow
> the admin to inspect the changes on the system compared to the rpm
> database. I have found similar ideas while searching for an
> existing solution
> (http://tomayko.com/writings/MinimalSystemBackups) but have not
> found an already working one :-) The script tries to be careful to 
> create only the files which are not changed (in comparison to the
> rpm db) also taking into account the changes that can occur due to
> prelinking. Of course a full dar backup would be more suitable for
> restoring, since rebuilding would certainly take more time (plus
> admin would have to be careful to install the same version of rpm
> packages so that the stored config files would not overwrite a
> possibly newer config format in a newer rpm package). But such a
> diff backup would be easier to inspect and thus verify the validity
> of a full backup made at the same time.
> 
> A more advanced backup script could use this approach to create a
> chain of incremental backups, each storing only the subset of
> changes compared to the rpm database and the previous diff backup
> (which could be made against a previous state of the rpm database)
> by merging a new rpm snapshot backup with the previous diff backup
> (one note, I have just noticed in the 2.4.10 man page that merging
> of isolated catalogues is not supported - I have tried searching
> for why this has changed, but I only found an explanation that "If 
> you merge two extracted catalogues you will get an extracted
> catalogue").

The reason is that since 2.4.0 an isolated catalogue is not more a
lighter version if the internal catalogue, in particular it contains
the offsets in the archive where to find data and EA and is thus able
to be used in place of the internal catalogue at restoration time (-x
<archive> -A <isolated-cat>).

To avoid mistakes if one tries to restore an archive using an isolated
catalogue from another backup each archive has an internal label
(randomly chosen and based on creation time) that sticks with the data
set. When isolating a catalogue this label is kept equal to the one of
the original backup. Slicing using dar_xform does not change that label,
but merging does and thus the isolated catalogue of an archive used as
source for merging cannot be used with the merged archive.

As you guess merging two isolated catalogues would loose the ability
of the resulting archive to backup the internal catalogue of any
backup, it has thus been forbidden to rather force the creation of an
isolated catalogue of the merged archive.

Note that a snapshot backup is no more the same as a catalogue since
release 2.4.0. You can merge snapshots backups at will with full,
differential backup or another snapshot.

At merging time, you may find useful to tune the resulting archive by
playing with the overwriting policy (-/ option) eventually you may
find interesting to proceed to decremental backups (-ad option). see
/etc/darrc there are some examples of use like the full-from-diff
target (I can comment it out if you need, as it may not sound very
simple and clear to read).

> 
> Does this make any sense or have I just wasted a perfectly good
> Sunday afternoon?

I think it *does* make sens. While OK, I am not sure to have very
deeply understood your script but that's because I'm not yet
python-ready ;-)

> 
> Regards, gulikoza
> 
> 

Regards,
Denis.


-----BEGIN PGP SIGNATURE-----
Version: GnuPG v1.4.12 (GNU/Linux)
Comment: Using GnuPG with Mozilla - http://enigmail.mozdev.org/

iQIVAwUBUua/FQgxsL0D2LGCAQJdmQ//WnAmGQR5lbGEoVR5qMzQBo8+D7aRuEiB
IYoBNx4cvNXNrU3dUblEz68/XzfY+cs8pAjyQAFM9CamqU9gbRI4i5DKjf0+fw4U
khVTw/Tnd5wbQzSZTSBUxG+yvbj14FjxBXxuosVhG4erzAZ2lQhHbmYpbnnu9Ks7
phXuBD4l2GKUgWcBpQA4538P1EVCBZwg/l3zY+DN1HPrUE4g9i3Z2j52bUxr7uVH
p8G9pfDCSJqFrawFVoWPuI2K1xCMBuhpePapusgj38a2eMDLNUOazqcQmnLGxO6a
OoeTkVhJZLSVNG7abGP0ejb12kwNW338fuLQnNEsB+fNdm51S/cpEkz3oHpbFbMA
PJ59xKZFr+zIq4Kie/Lhw2qqQOZEUl79FmqKVWcqB8648fwPmF921kknhFETJ4vJ
C5fYzGTQ5rymeL57N7Lavm0q/r7FHKrlRFYuZ8Ksar9MWar4LVOOeK0zds4XJEPz
dmBkZyWFmEAGX9bm0LunckmaIJ0oIBHDSAdp/RKKmNxz3s1rqul0aOTRUdwow4TL
EWHe9C0VQue3Sn1BkV96lthLX8l+hVAAreaySLB6x/L48AXi9rAgHbY5SlBnCTpH
NUoAldoK0QWEhP1iRTdGmmkOn0y4jqQOwmZkyJ5J2s8DVLsMPC10BFmCKWqs46eX
wEN+8pyKE08=
=oOoL
-----END PGP SIGNATURE-----

------------------------------------------------------------------------------
CenturyLink Cloud: The Leader in Enterprise Cloud Services.
Learn Why More Businesses Are Choosing CenturyLink Cloud For
Critical Workloads, Development Environments & Everything In Between.
Get a Quote or Start a Free Trial Today. 
http://pubads.g.doubleclick.net/gampad/clk?id=119420431&iu=/4140/ostg.clktrk