Re: Over 1.5GB file limit?
Pat Thoyts <[email protected]> Tue, 29 Mar 2011 18:57:15 +0100
| Newsgroups | gmane.comp.db.metakit |
|---|---|
| Message-ID | <[email protected]> |
--00151773e3505f3396049fa2ca9e Content-Type: text/plain; charset=ISO-8859-1 On 29 March 2011 18:21, Thadeus Burgess <[email protected]> wrote: > 3rd tries a charm to post this message to the group ! > > It is glad to know metakit is being actively used. > > We are using it to log data that comes in every second from many nodes, all > time stamped and need to be able to store potentially millions of rows. We > do plan on running metakit on some cloud servers, however this also has to > run on some Atom class processor typically found in netbooks (since this is > where the data is coming from). > > Currently we are using PostgreSQL in production, yet unfortunately it is > VERY slow and is unable to keep up with the data demand. Yes Postgres is > configured, has proper indexing etc etc.... > > In preliminary tests... > > We can stuff about 11million records into 2GB. There shouldn't be more than > 2million records in a given day. Using a time-series based file partitioning > (year -> month -> day -> data.mk) the 2GB limit should not occur. > > Performing a select on a range in the 11M records takes about .12 seconds > (SQL -> 15 minutes) > Performing the above select and sorting it by another column .40 seconds > (SQL -> 19 minutes) > Performing an avg on a column in .96 seconds (SQL -> 17 minutes) > Finding min/max on a column 1 second (SQL -> 23 minutes... yeah I don't get > that either) > > It may not be the fastest out there, but its still trillions of times faster > than any SQL would ever be. > > I have read in certain places that HDF5 outperforms metakit. However HDF5 > looks more complicated to get installed and running. > > We will also continue to be using SQL for our metadata, since the metadata > is relational a relational database makes sense there. We store spectroscopic data into metakit files. We've been doing so for about 8 years now and last time I counted up the local datafiles we have nearly 0.75TB stored in these things. We store each individial spectrum as a row of some values and a blob of binary data. We found that over some threshold (10 000 rows or so) the default configuration doesnt scale too well. If you intend to have lots of rows, ensure you use blocked views. Speed wise - you will have to test your setup. I'm now moving away from our metakit based files so that we can memory map the data directly as both the 2GB limit and the time taken to extract an individual dataset are now limiting. (The time issue is not metakit's fault - we also serialized COM objects into the data stream - in retrospect this was a mistake). If your data is constant sized, then fastest and least resource intensive is to simply append blocks onto the end of the file when writing and use memory mapped I/O with a suitable window size when reading them later. Anyway: I include a sample of blocked view usage I made today while trying to see if I could create a >2GB file on my 64bit Windows7 system (using a 64bit exe). It turns out I cannot. It is possible this is a limit in the c4_FileStrategy which is where metakit abstracts the O/S part of the file handling. Once this test exceeds 2GB it no longer creates the file (bigdemo -create 250000). Under 2GB its ok and bigdemo -create 100000 & bigdemo -check 80000 was working nice and seemingly fast. Pat Thoyts -- You received this message because you are subscribed to the "metakit" group. To post to this group, send email to [email protected] To unsubscribe from this group, send email to [email protected] For more options, visit this group at http://groups.google.com/group/metakit?hl=en --00151773e3505f3396049fa2ca9e Content-Type: text/x-c++src; charset=US-ASCII; name="bigdemo.cpp" Content-Disposition: attachment; filename="bigdemo.cpp" Content-Transfer-Encoding: base64 X-Attachment-Id: f_glv4n58h0 I2luY2x1ZGUgPHN0ZGlvLmg+DQojaW5jbHVkZSA8c3RkbGliLmg+DQojaW5jbHVkZSA8c3RyaW5n Lmg+DQojaW5jbHVkZSA8bWs0Lmg+DQoNCnN0YXRpYyB2b2lkDQpyYW5kZmlsbChjaGFyICpkYXRh LCBzaXplX3QgbGVuKQ0Kew0KICAgIGZvciAoc2l6ZV90IG4gPSAwOyBuIDwgbGVuOyArK24pDQog ICAgICAgICpkYXRhKysgPSBjaGFyKHJhbmQoKSAqIDI2ICsgOTUpOw0KfQ0KDQppbnQNCm1haW4o aW50IGFyZ2MsIGNoYXIgKmFyZ3ZbXSkNCnsNCiAgICBpZiAoYXJnYyAhPSAzKSB7DQogICAgICAg IHByaW50ZigidXNhZ2U6IGJpZ2RlbW8gLWNyZWF0ZXwtY2hlY2sgY291bnRcbiIpOw0KICAgICAg ICByZXR1cm4gMDsNCiAgICB9DQoNCiAgICB1bnNpZ25lZCBsb25nIGNvdW50ID0gc3RydG91bChh cmd2WzJdLCBOVUxMLCAwKTsNCiAgICANCiAgICBjNF9TdG9yYWdlIHN0b3JhZ2UoImJpZ2RlbW8u ZGF0Iix0cnVlKTsNCiAgICBjNF9WaWV3IHZSb290ID0gc3RvcmFnZS5HZXRBcygic3R1ZmZbX0Jb bmFtZTpTLGRhdGE6Ql1dIik7DQogICAgYzRfVmlldyB2U3R1ZmYgPSB2Um9vdC5CbG9ja2VkKCku T3JkZXJlZCgpOw0KICAgIGM0X1N0cmluZ1Byb3Agck5hbWUoIm5hbWUiKTsNCiAgICBjNF9CeXRl c1Byb3AgckRhdGEoImRhdGEiKTsNCg0KICAgICBpZiAoc3RyY21wKCItY3JlYXRlIiwgYXJndlsx XSkgPT0gMCkgew0KICAgIA0KICAgICAgICBjaGFyICpjaGFycyA9IG5ldyBjaGFyWzEwMjQwMDBd Ow0KICAgICAgICByYW5kZmlsbChjaGFycywgMTAyNDApOw0KICAgICAgICBjNF9CeXRlcyBkYXRh KGNoYXJzLCAxMDI0MCk7DQogICAgICAgIA0KICAgICAgICBmb3IgKHNpemVfdCBuID0gMDsgbiA8 IGNvdW50OyArK24pIHsNCiAgICAgICAgICAgIGNoYXIgbmFtZVszNF07DQogICAgICAgICAgICBz cHJpbnRmX3MobmFtZSwgc2l6ZW9mKG5hbWUpLCAiaXRlbSVsdSIsIG4pOw0KICAgICAgICAgICAg Y2hhcnNbMF0gPSBjaGFyKG4pOw0KICAgICAgICAgICAgdlN0dWZmLkFkZChyTmFtZVtuYW1lXSAr IHJEYXRhW2RhdGFdKTsNCiAgICAgICAgfQ0KICAgICAgICANCiAgICAgICAgc3RvcmFnZS5Db21t aXQoKTsNCiAgICB9IGVsc2Ugew0KICAgICAgICBjaGFyIG5hbWVbMzRdOw0KICAgICAgICBzcHJp bnRmX3MobmFtZSwgc2l6ZW9mKG5hbWUpLCAiaXRlbSVsdSIsIGNvdW50KTsNCiAgICAgICAgDQog ICAgICAgIGludCByb3cgPSB2U3R1ZmYuRmluZChyTmFtZVtuYW1lXSk7DQogICAgICAgIHByaW50 ZigibG9va2luZyBmb3IgJyVzJyByZXR1cm5lZCByb3cgJWRcbiIsIG5hbWUsIHJvdyk7DQogICAg ICAgIGlmIChyb3cgPT0gLTEpDQogICAgICAgICAgICBwcmludGYoImVycm9yXG4iKTsNCiAgICAg ICAgZWxzZQ0KICAgICAgICAgICAgcHJpbnRmKCJmb3VuZCAnJXMnXG4iLCAoY29uc3QgY2hhciAq KXJOYW1lKHZTdHVmZltyb3ddKSk7DQogICAgfQ0KICAgIHJldHVybiAwOw0KfQ0K --00151773e3505f3396049fa2ca9e--