Re: Blobs and data science, memory mapping and efficient appends

Jim Fulton <[email protected]> Sat, 21 Apr 2018 12:12:48 -0600
Newsgroups gmane.comp.web.zope.zodb
Message-ID <CAPDm-Fg_yTJwJ45BhksK+=Yikd2cD1T4u+6GvpZytGFEG0K2+w@mail.gmail.com>
Ha, right after sending this I saw this:

https://twitter.com/garybernhardt/status/987447046150934528

Jim

On Sat, Apr 21, 2018 at 12:09 PM, Jim Fulton <[email protected]> wrote:

> I've been working on data management to support data science and think a
> lot about how ZODB can help.  For data analysis in Python, Python objects
> and thus ZODB objects aren't suitably efficient. People typically use dense
> formats, such as numpy and related data formats.  On-disk formats for this
> data can often be memory mapped to great profit, allowing analysis of data
> larger than will fit in RAM at once.  Doing this is straightforward with
> ZODB blobs, thanks to the fact that blobs are exposed as Python files. I've
> done this. It's easy and cool. There is much opportunity here.
>
> I'm thinking of how to efficiently handle new data, which often takes the
> form of appends.  Think of use cases where we always add data to a dataset
> without updating existing data.  Normally, when we open a blob for append,
> we make a copy of committed data, and when we commit in
> non-shared-directory mode, we copy the entire file to the server.  If files
> are large, but appends are small, this is wasteful, prohibitively so for
> some applications.
>
> I've been pondering extending blobs to allow multiple versions to share a
> single blob file when subsequent versions only append data.  The blob
> object we have now would be modified to store a blob size. Different
> versions sharing a single file, would have different sizes.  There are a
> number of details to work out, but I think they can be dealt with fairly
> straightforwardly.
>
> Advantages include:
>
>    - Don't need to copy when opening for append.
>    - Only need to send new data to the server on commit.
>    - Only need to fetch new data from the server when loading.
>    - There might be some interesting opportunities for conflict
>    resolution, allow simultaneous append.
>
> Thoughts?
>
> Is anyone interested in working on this with me?
>
> I'm thinking of hacking on this at PyCon.
>
> Jim
>
> P.S. See also https://github.com/Nexedi/wendelin.core which is using ZODB
> for data science.
>
> --
> Jim Fulton
> http://jimfulton.info
>



-- 
Jim Fulton
http://jimfulton.info

-- 
You received this message because you are subscribed to the Google Groups "zodb" group.
To unsubscribe from this group and stop receiving emails from it, send an email to [email protected].
For more options, visit https://groups.google.com/d/optout.