Re: Packing, again
Richard Waid <[email protected]>
| Newsgroups | gmane.comp.web.zope.zodb.dirstorage |
|---|---|
| Organization | IOPEN Technologies Ltd. |
| Message-ID | <[email protected]> |
On Sat, 2004-09-04 at 15:13 +0100, Toby Dickenson wrote: > On Saturday 04 Sep 2004 11:57, Richard Waid wrote: > > > I'm currently exporting the entire database to a file (via the > > export mechanism in ZODB) to mitigate against any possible dataloss. > > I recommend doing this to a different filesystem on a different disk > controller, to guard against hardware problems. Ideally, please take a tar > backup of both replica and master storage directories. I normally do. It does take an extremely long time to tar a 8-10gig storage consisting of lots of small files :) > > I've just tried a pack of our largest dirstorage, > > You mentioned you have been trying the 'Minimal' packing method. Was that in > use here, or with this storage in the past? Yes, it was with 'Minimal', which BTW, seemed to be working a _lot_ quicker than with the 'permission bit' method. I've never used it on this storage in the past. > > Perhaps more importantly, what can be done to fix it (without data > > loss)? > > If it does look like packing has gone wrong, then it is definitely safe to > repair a storage by 'undeleting' files by renaming them to remove the > -XXXXXX-deleted suffix or copying individual files from the replica. This is > completely safe provided you do not overwrite any file. (overwriting a > damaged file might be the right way to fix a problem, but you need to take > care) > > $ python dumpdsf.py ~/projects/Zope/var/ds/A/x/packed > /home/toby/projects/Zope/var/ds/A/x/packed > current rev 03577F55D34D77CC > transaction timestamp Mon Aug 30 16:17:49 2004 I'm getting: current rev 035773A106845322 transaction timestamp Sun Aug 29 01:21:01 2004 Quite right. What I _think_ happened was that I packed the database, but I had to abort it (basically the load meant that we had a cascading load problem -- requests queued until it locked solid). I, perhaps stupidly in hindsight, was running the database with delay_delete: 0 (the comment said 'this is appropriate if you are using a stable version of DirectoryStorage'). I've changed this now because obviously I can actually delete the files at my leisure and use that to reduce the length of the load on the system. I can't find any 'deleted' transactions, which I'm assuming (and the code backs this up) means that if you have 'delay_delete: 0' the transactions are deleted in a first sweep. I should have checked the code first ... I guess I assumed that if you had delay_delete:0 it would rename then perform another sweep to remove. So... what does this mean? I'm assuming the missing files were removed in the aborted pack. I'm also assuming that if they were deleted they probably weren't needed anyway (which explains while our application hasn't turned into a steaming mess yet :)). I'm guessing I could fix this by copying the transactions from the replica, though that could be pretty tedious. What I've learnt: Don't use delay_delete: 0, especially if you think you might have to abort :) --Richard ------------------------------------------------------- This SF.Net email is sponsored by BEA Weblogic Workshop FREE Java Enterprise J2EE developer tools! Get your free copy of BEA WebLogic Workshop 8.1 today. http://ads.osdn.com/?ad_id=5047&alloc_id=10808&op=click