Re: Over 1.5GB file limit?

Pat Thoyts <[email protected]> Tue, 29 Mar 2011 18:57:15 +0100
Newsgroups gmane.comp.db.metakit
Message-ID <[email protected]>
--00151773e3505f3396049fa2ca9e
Content-Type: text/plain; charset=ISO-8859-1

On 29 March 2011 18:21, Thadeus Burgess <[email protected]> wrote:
> 3rd tries a charm to post this message to the group !
>
> It is glad to know metakit is being actively used.
>
> We are using it to log data that comes in every second from many nodes, all
> time stamped and need to be able to store potentially millions of rows. We
> do plan on running metakit on some cloud servers, however this also has to
> run on some Atom class processor typically found in netbooks (since this is
> where the data is coming from).
>
> Currently we are using PostgreSQL in production, yet unfortunately it is
> VERY slow and is unable to keep up with the data demand. Yes Postgres is
> configured, has proper indexing etc etc....
>
> In preliminary tests...
>
> We can stuff about 11million records into 2GB. There shouldn't be more than
> 2million records in a given day. Using a time-series based file partitioning
> (year -> month -> day -> data.mk) the 2GB limit should not occur.
>
> Performing a select on a range in the 11M records takes about .12 seconds
> (SQL -> 15 minutes)
> Performing the above select and sorting it by another column .40 seconds
> (SQL -> 19 minutes)
> Performing an avg on a column in .96 seconds (SQL -> 17 minutes)
> Finding min/max on a column 1 second (SQL -> 23 minutes... yeah I don't get
> that either)
>
> It may not be the fastest out there, but its still trillions of times faster
> than any SQL would ever be.
>
> I have read in certain places that HDF5 outperforms metakit. However HDF5
> looks more complicated to get installed and running.
>
> We will also continue to be using SQL for our metadata, since the metadata
> is relational a relational database makes sense there.

We store spectroscopic data into metakit files. We've been doing so
for about 8 years now and last time I counted up the local datafiles
we have nearly 0.75TB stored in these things. We store each individial
spectrum as a row of some values and a blob of binary data. We found
that over some threshold (10 000 rows or so) the default configuration
doesnt scale too well. If you intend to have lots of rows, ensure you
use blocked views.

Speed wise - you will have to test your setup. I'm now moving away
from our metakit based files so that we can memory map the data
directly as both the 2GB limit and the time taken to extract an
individual dataset are now limiting. (The time issue is not metakit's
fault - we also serialized COM objects into the data stream - in
retrospect this was a mistake). If your data is constant sized, then
fastest and least resource intensive is to simply append blocks onto
the end of the file when writing and use memory mapped I/O with a
suitable window size when reading them later.

Anyway: I include a sample of blocked view usage I made today while
trying to see if I could create a >2GB file on my 64bit Windows7
system (using a 64bit exe). It turns out I cannot. It is possible this
is a limit in the c4_FileStrategy which is where metakit abstracts the
O/S part of the file handling. Once this test exceeds 2GB it no longer
creates the file (bigdemo -create 250000). Under 2GB its ok and
bigdemo -create 100000 & bigdemo -check 80000 was working nice and
seemingly fast.

Pat Thoyts

-- 
You received this message because you are subscribed to the "metakit" group.
To post to this group, send email to [email protected]
To unsubscribe from this group, send email to [email protected]
For more options, visit this group at http://groups.google.com/group/metakit?hl=en
--00151773e3505f3396049fa2ca9e
Content-Type: text/x-c++src; charset=US-ASCII; name="bigdemo.cpp"
Content-Disposition: attachment; filename="bigdemo.cpp"
Content-Transfer-Encoding: base64
X-Attachment-Id: f_glv4n58h0

I2luY2x1ZGUgPHN0ZGlvLmg+DQojaW5jbHVkZSA8c3RkbGliLmg+DQojaW5jbHVkZSA8c3RyaW5n
Lmg+DQojaW5jbHVkZSA8bWs0Lmg+DQoNCnN0YXRpYyB2b2lkDQpyYW5kZmlsbChjaGFyICpkYXRh
LCBzaXplX3QgbGVuKQ0Kew0KICAgIGZvciAoc2l6ZV90IG4gPSAwOyBuIDwgbGVuOyArK24pDQog
ICAgICAgICpkYXRhKysgPSBjaGFyKHJhbmQoKSAqIDI2ICsgOTUpOw0KfQ0KDQppbnQNCm1haW4o
aW50IGFyZ2MsIGNoYXIgKmFyZ3ZbXSkNCnsNCiAgICBpZiAoYXJnYyAhPSAzKSB7DQogICAgICAg
IHByaW50ZigidXNhZ2U6IGJpZ2RlbW8gLWNyZWF0ZXwtY2hlY2sgY291bnRcbiIpOw0KICAgICAg
ICByZXR1cm4gMDsNCiAgICB9DQoNCiAgICB1bnNpZ25lZCBsb25nIGNvdW50ID0gc3RydG91bChh
cmd2WzJdLCBOVUxMLCAwKTsNCiAgICANCiAgICBjNF9TdG9yYWdlIHN0b3JhZ2UoImJpZ2RlbW8u
ZGF0Iix0cnVlKTsNCiAgICBjNF9WaWV3IHZSb290ID0gc3RvcmFnZS5HZXRBcygic3R1ZmZbX0Jb
bmFtZTpTLGRhdGE6Ql1dIik7DQogICAgYzRfVmlldyB2U3R1ZmYgPSB2Um9vdC5CbG9ja2VkKCku
T3JkZXJlZCgpOw0KICAgIGM0X1N0cmluZ1Byb3Agck5hbWUoIm5hbWUiKTsNCiAgICBjNF9CeXRl
c1Byb3AgckRhdGEoImRhdGEiKTsNCg0KICAgICBpZiAoc3RyY21wKCItY3JlYXRlIiwgYXJndlsx
XSkgPT0gMCkgew0KICAgIA0KICAgICAgICBjaGFyICpjaGFycyA9IG5ldyBjaGFyWzEwMjQwMDBd
Ow0KICAgICAgICByYW5kZmlsbChjaGFycywgMTAyNDApOw0KICAgICAgICBjNF9CeXRlcyBkYXRh
KGNoYXJzLCAxMDI0MCk7DQogICAgICAgIA0KICAgICAgICBmb3IgKHNpemVfdCBuID0gMDsgbiA8
IGNvdW50OyArK24pIHsNCiAgICAgICAgICAgIGNoYXIgbmFtZVszNF07DQogICAgICAgICAgICBz
cHJpbnRmX3MobmFtZSwgc2l6ZW9mKG5hbWUpLCAiaXRlbSVsdSIsIG4pOw0KICAgICAgICAgICAg
Y2hhcnNbMF0gPSBjaGFyKG4pOw0KICAgICAgICAgICAgdlN0dWZmLkFkZChyTmFtZVtuYW1lXSAr
IHJEYXRhW2RhdGFdKTsNCiAgICAgICAgfQ0KICAgICAgICANCiAgICAgICAgc3RvcmFnZS5Db21t
aXQoKTsNCiAgICB9IGVsc2Ugew0KICAgICAgICBjaGFyIG5hbWVbMzRdOw0KICAgICAgICBzcHJp
bnRmX3MobmFtZSwgc2l6ZW9mKG5hbWUpLCAiaXRlbSVsdSIsIGNvdW50KTsNCiAgICAgICAgDQog
ICAgICAgIGludCByb3cgPSB2U3R1ZmYuRmluZChyTmFtZVtuYW1lXSk7DQogICAgICAgIHByaW50
ZigibG9va2luZyBmb3IgJyVzJyByZXR1cm5lZCByb3cgJWRcbiIsIG5hbWUsIHJvdyk7DQogICAg
ICAgIGlmIChyb3cgPT0gLTEpDQogICAgICAgICAgICBwcmludGYoImVycm9yXG4iKTsNCiAgICAg
ICAgZWxzZQ0KICAgICAgICAgICAgcHJpbnRmKCJmb3VuZCAnJXMnXG4iLCAoY29uc3QgY2hhciAq
KXJOYW1lKHZTdHVmZltyb3ddKSk7DQogICAgfQ0KICAgIHJldHVybiAwOw0KfQ0K
--00151773e3505f3396049fa2ca9e--