Re: Loading Resources From a Python Module/Package

Brett Cannon <[email protected]> Sat, 31 Jan 2015 21:22:42 +0000
Newsgroups gmane.comp.python.import
Message-ID <CAP1=2W6kg2tWbS_1O-BTSY-L5=AQ42Gmm8n1RRdYxT_63j1PGA@mail.gmail.com>
On Sat Jan 31 2015 at 12:28:07 PM Donald Stufft <[email protected]> wrote:

> On Jan 31, 2015, at 12:00 PM, Brett Cannon <[email protected]> wrote:
>
>
>
> On Sat Jan 31 2015 at 11:43:55 AM Donald Stufft <[email protected]> wrote:
>
>> On Jan 31, 2015, at 11:31 AM, Brett Cannon <[email protected]> wrote:
>>
>>
>>
>> On Sat Jan 31 2015 at 10:54:22 AM Paul Moore <[email protected]> wrote:
>>
>>> On 31 January 2015 at 15:47, Donald Stufft <[email protected]> wrote:
>>> >> It's certainly possible to add a new API that loads resources based on
>>> >> a relative name, but you'd have to specify relative to *what*.
>>> >> get_data explicitly ducks out of making that decision.
>>> >
>>> > data = __loader__.get_bytes(__name__, “logo.gif”)
>>>
>>> Quite possibly. It needs a bit of fleshing out to make sure it doesn't
>>> prohibit sharing of loaders, etc, in the way Brett mentions.
>>
>>
>> By specifying the package anchor point I don't think it does.
>>
>>
>>> Also, the
>>> fact that it needs __name__ in there feels wrong - a bit like the old
>>> version of super() needing to be told which class it was being called
>>> from.
>>
>>
>> You can't avoid that. This is the entire reason why loader reuse is a
>> pain; you **have** to specify what to work off of, else its ambiguous and a
>> specific feature of a specific loader.
>>
>> But this is only an issue when you are trying to access a file relative
>> to the package/module you're in. Otherwise you're going to be specifying a
>> string constant like 'foo.bar'.
>>
>>
>>> But in principle I don't object to finding a suitable form of
>>> this.
>>>
>>> And I like the name get_bytes - much more explicit in these Python 3
>>> days of explicit str/bytes distinctions :-)
>>
>>
>> One unfortunate side-effect from having a new method to return bytes from
>> a data file is that it makes get_data() somewhat redundant. If we make it
>> get_data_filename(package_name, path) then it can return an absolute path
>> which can then be passed to get_data() to read the actual bytes. If we
>> create importlib.resources as Donald has suggested then all of this can be
>> hidden  behind a function and users don't have to care about any of this,
>> e.g. importlib.resources.read_data(module_anchor, path).
>>
>>
>> I think we actually have to go the other way, because only some Loaders
>> will be able to actually return a filename (returning a filename is
>> basically an optimization to prevent needing to call get_data and write
>> that out to a temporary directory) but pretty much any loader should
>> theoretically be able to support get_data.
>>
>
> Why can only some loaders return a filename? As I have said, loaders can
> return an opaque string to simulate a path if necessary.
>
>
> Because the idea behind get_data_filename() is that it returns a path that
> can be used regularly by APIs that expect to be handed a file on the file
> system.
>

In my head that expectation is not placed on the method.


> Simulating a path with an opaque string isn’t good enough because, for
> example, OpenSSL doesn’t know how to open /data/foo.zip/foobar/cacert.pem.
> The idea here is that _if_ a regular file system path is available for a
> particular resource file then Loader().get_data_filename() would return it,
> otherwise it’d return None (or not exist at all).
>
> This means that pkgutil.get_data_filename (or
> importlib.resources.get_filename) can attempt to call
> Loader().get_data_filename() and just return that path if one exists on the
> file system already, and if it doesn’t then it can create a temporary file
> and call Loader.get_data() and write the data to that temporary file and
> return the path to that.
>

See I'm not even attempting to guarantee there is any API that will return
a reasonable file system path as the import API makes no such guarantees.
If an API like OpenSSL requires a file on the filesystem then you will have
to write to a temporary file and that's just life. That's the same as if
everything was stored in a zip file anyway.


>
>
>
>>
>> I think it is redundant but given that it’s a new API (passing module and
>> a “resource path”) I think it makes sense. The old get_data API can be
>> deprecated but left in for compatibility reasons if we want (sort of like
>> Loader().load_module() -> Loader().exec_module()).
>>
>
> If we do that then there would have to be a way to specify how to read the
> bytes for the module code itself since get_data() is used in the
> implementation of import by coupling it with get_filename() (which is why
> I'm trying not have to drop get_filename()/get_data() and instead come up
> with some new approach to reading bytes since the current approach is very
> composable). So get_bytes() would need a way to signal that you don't want
> some data file but the bytes for the module. Maybe if the path section is
> unspecified then that's a signal that the module's bytes is wanted and not
> some data file?
>
>
> Perhaps trying to read modules and resource files with the same method is
> the wrong approach?
>

If we are going to do that then we might as well deprecate all the methods
that try to expose reading data and paths as the PEP 302 APIs tried to
expose it uniformly.


>
> Maybe instead we should do: https://bpaste.net/show/b25b7e8dc8f0
>

That seems like a bit much, e.g. why do you needs bytes **and** and a
file-like object() when you get the former from the latter? And why do you
need the path argument when you can get the path off the file-like object
if it's an actual file object?

-Brett


>
> This means that we’re not talking about “data” files, but “resource”
> files. This also removes the idea that you can call Loader.set_data() on
> those files (like i’ve seen in the implementation).
>
>
>
>>
>>
>> One thing to consider is do we want to allow anything other than
>> filenames for the path part? Thanks to namespace packages every directory
>> is essentially a package, so we could say that the package anchor has to
>> encapsulate the directory and the path bit can only be a filename. That
>> gets us even farther away from having the concept of file paths being
>> manipulated in relation to import-related APIs.
>>
>>
>> I think we do want to allow directories, it’s not unusual to have
>> something like:
>>
>> warehouse
>> ├── __init__.py
>> ├── templates
>> │   ├── accounts
>> │   │   └── profile.html
>> │   └── hello.html
>> ├── utils
>> │   └── mapper.py
>> └── wsgi.py
>>
>> Conceptually templates isn’t a package (even though with namespace
>> packages it kinda is) and I’d want to load profile.html by doing something
>> like:
>>
>> importlib.resources.get_bytes(“warehouse”,
>> “templates/accounts/profile.html”)
>>
>
> Where I would be fine with get_bytes('warehouse.templates.accounts',
> 'profile.html')  =)
>
>
>>
>> In pkg_resources the second argument to that function is a “resource
>> path” which is defined as a relative to the given module/package and it
>> must use / to denote them. It explicitly says it’s not a file system path
>> but a resource path. It may translate to a file system path (as is the case
>> with the FileLoader) but it also may not (as is the case with a theoretical
>> S3Loader or PostgreSQLLoader).
>>
>
> Yep, which is why I'm making sure if we have paths we minimize them as
> they instantly make these alternative loader concepts a bigger pain to
> implement.
>
>
>> How you turn a warehouse + a resource path into some data (or whatever
>> other function we support) is an implementation detail of the Loader.
>>
>>
>> And just so I don't forget it, I keep wanting to pass an actual module in
>> so the code can extract the name that way, but that prevents the __name__
>> trick as you would have to import yourself or grab the module from
>> sys.modules.
>>
>>
>> Is an actual module what gets passed into Loader().exec_module()?
>>
>
> Yes.
>
>
>> If so I think it’s fine to pass that into the new Loader() functions and
>> a new top level API in importlib.resources can do the things needed to turn
>> a string into a module object. So instead of doing
>> __loader__.get_bytes(__name__, “logo.gif”) you’d do
>> importlib.resources.get_bytes(__name__, “logo.gif”).
>>
>
> If we go the route of importlib.resources then that seems like a
> reasonable idea, although we will need to think through the ramifications
> to exec_module() itself although I don't think there were be any issues.
>
> And if we do go with importlib.resources I will probably want to make it
> available on PyPI with appropriate imp/pkgutil fallbacks to help people
> transitioning from Python 2 to 3.
>
> ---
> Donald Stufft
> PGP: 7C6B 7C5D 5E2B 6356 A926 F04F 6E3C BCE9 3372 DCFA
>
>

_______________________________________________
Import-SIG mailing list
[email protected]
https://mail.python.org/mailman/listinfo/import-sig