Re: Create a catalog from the rpm database
Denis Corbin <[email protected]> Mon, 27 Jan 2014 21:18:29 +0100
| Newsgroups | gmane.comp.sysutils.backup.dar.general |
|---|---|
| Message-ID | <[email protected]> |
-----BEGIN PGP SIGNED MESSAGE----- Hash: SHA1 On 27/01/2014 09:19, gulikoza wrote: > Hello, Hello, > > I use dar extensively for backups, usually separating system (the > linux itself) and content backups. :) > I store system backups also in my archive in order to have some > history of the changed configurations and such. Do you mean that the system backups take place inside the content backup? Not sure to understand (... sorry its late here :@) ) to have the history of changes you might be using dar_manager, unless I don't exactly what you need to do... > > The problem is that this is becoming a big organizational > nightmare. A simple update might change a lot of files, bloating > backup size Yes, I've noted that upgrading a package (.deb or .rpm) often reinstall 90% of the file concerned by that package ... which dar detects as changed and save them again. You can eventually ignore some attributes modifications, like permission or ownership while the file size is the same [--comparison-field option] but at the risk you miss some updated files... > and if you multiply that with a few dozen servers, checking each > and every backup is quite a task. Another possibility would be to upgrade a reference server with packaging tools and take differential backup of it with dar that you push to other servers, no? > However, most of the stuff is rarely needed as files not > specifically changed are better tracked by rpm itself (having a > list of installed packages + changes would allow you to rebuild the > same system). I have thought of creating a script that would list > rpm changed and untracked files and include only those in the > system backup, but there are some exceptions. Some files might be > changed on the disk, but not listed in rpm verify for instance the > rpm database itself: > > %verify(not md5 size mtime) %ghost %config(missingok,noreplace) > /var/lib/rpm/* > > The files in /var/lib/rpm will not be shown as modified by rpm -V, > but will be included in rpm -ql. Any script working only with the > list of changed and untracked files, would fail to include these in > the backup. > > At first I thought of somehow interfacing dar directly with the rpm > database but then I decided for a simpler approach. The attached > script tries to create an empty root tree (sparse files) from the > rpm database that can be used to create a dar catalogue (-A +, a > snapshot backup) which in turn can be used as a reference for a > system backup. Such a backup would be pretty lightweight and allow > the admin to inspect the changes on the system compared to the rpm > database. I have found similar ideas while searching for an > existing solution > (http://tomayko.com/writings/MinimalSystemBackups) but have not > found an already working one :-) The script tries to be careful to > create only the files which are not changed (in comparison to the > rpm db) also taking into account the changes that can occur due to > prelinking. Of course a full dar backup would be more suitable for > restoring, since rebuilding would certainly take more time (plus > admin would have to be careful to install the same version of rpm > packages so that the stored config files would not overwrite a > possibly newer config format in a newer rpm package). But such a > diff backup would be easier to inspect and thus verify the validity > of a full backup made at the same time. > > A more advanced backup script could use this approach to create a > chain of incremental backups, each storing only the subset of > changes compared to the rpm database and the previous diff backup > (which could be made against a previous state of the rpm database) > by merging a new rpm snapshot backup with the previous diff backup > (one note, I have just noticed in the 2.4.10 man page that merging > of isolated catalogues is not supported - I have tried searching > for why this has changed, but I only found an explanation that "If > you merge two extracted catalogues you will get an extracted > catalogue"). The reason is that since 2.4.0 an isolated catalogue is not more a lighter version if the internal catalogue, in particular it contains the offsets in the archive where to find data and EA and is thus able to be used in place of the internal catalogue at restoration time (-x <archive> -A <isolated-cat>). To avoid mistakes if one tries to restore an archive using an isolated catalogue from another backup each archive has an internal label (randomly chosen and based on creation time) that sticks with the data set. When isolating a catalogue this label is kept equal to the one of the original backup. Slicing using dar_xform does not change that label, but merging does and thus the isolated catalogue of an archive used as source for merging cannot be used with the merged archive. As you guess merging two isolated catalogues would loose the ability of the resulting archive to backup the internal catalogue of any backup, it has thus been forbidden to rather force the creation of an isolated catalogue of the merged archive. Note that a snapshot backup is no more the same as a catalogue since release 2.4.0. You can merge snapshots backups at will with full, differential backup or another snapshot. At merging time, you may find useful to tune the resulting archive by playing with the overwriting policy (-/ option) eventually you may find interesting to proceed to decremental backups (-ad option). see /etc/darrc there are some examples of use like the full-from-diff target (I can comment it out if you need, as it may not sound very simple and clear to read). > > Does this make any sense or have I just wasted a perfectly good > Sunday afternoon? I think it *does* make sens. While OK, I am not sure to have very deeply understood your script but that's because I'm not yet python-ready ;-) > > Regards, gulikoza > > Regards, Denis. -----BEGIN PGP SIGNATURE----- Version: GnuPG v1.4.12 (GNU/Linux) Comment: Using GnuPG with Mozilla - http://enigmail.mozdev.org/ iQIVAwUBUua/FQgxsL0D2LGCAQJdmQ//WnAmGQR5lbGEoVR5qMzQBo8+D7aRuEiB IYoBNx4cvNXNrU3dUblEz68/XzfY+cs8pAjyQAFM9CamqU9gbRI4i5DKjf0+fw4U khVTw/Tnd5wbQzSZTSBUxG+yvbj14FjxBXxuosVhG4erzAZ2lQhHbmYpbnnu9Ks7 phXuBD4l2GKUgWcBpQA4538P1EVCBZwg/l3zY+DN1HPrUE4g9i3Z2j52bUxr7uVH p8G9pfDCSJqFrawFVoWPuI2K1xCMBuhpePapusgj38a2eMDLNUOazqcQmnLGxO6a OoeTkVhJZLSVNG7abGP0ejb12kwNW338fuLQnNEsB+fNdm51S/cpEkz3oHpbFbMA PJ59xKZFr+zIq4Kie/Lhw2qqQOZEUl79FmqKVWcqB8648fwPmF921kknhFETJ4vJ C5fYzGTQ5rymeL57N7Lavm0q/r7FHKrlRFYuZ8Ksar9MWar4LVOOeK0zds4XJEPz dmBkZyWFmEAGX9bm0LunckmaIJ0oIBHDSAdp/RKKmNxz3s1rqul0aOTRUdwow4TL EWHe9C0VQue3Sn1BkV96lthLX8l+hVAAreaySLB6x/L48AXi9rAgHbY5SlBnCTpH NUoAldoK0QWEhP1iRTdGmmkOn0y4jqQOwmZkyJ5J2s8DVLsMPC10BFmCKWqs46eX wEN+8pyKE08= =oOoL -----END PGP SIGNATURE----- ------------------------------------------------------------------------------ CenturyLink Cloud: The Leader in Enterprise Cloud Services. Learn Why More Businesses Are Choosing CenturyLink Cloud For Critical Workloads, Development Environments & Everything In Between. Get a Quote or Start a Free Trial Today. http://pubads.g.doubleclick.net/gampad/clk?id=119420431&iu=/4140/ostg.clktrk