RE: site-wide metadata discovery
"Chad Everett" <yahoogroups-ipyKpt5L2sG2CP6mJM4Q19i2O/[email protected]>
| Newsgroups | gmane.network.syndication.discuss |
|---|---|
| Message-ID | <350411BE3B7FF4488CA7797C9F1E90A6E306D4@exchange.ibmi.informedbeverage.com> |
Sorry for thinking this through in the list. Every time I hit send, another question comes to mind. So here's the next one: Would a single "metadata index file" be the most useful implementation, or would it be better for multiple records, for each type of metadata/function? Consider: #Feed-Index: feedlist.xml #Index-Format: http://purl.org/ocs/directory/0.5/ #Allow-Recursion: true #Site-Index: contents.html #Index-Format: http://www.w3.org/TR/xhtml11/ #Allow-Recursion: true #Site-Content: linkinfo #FOAF: foaf.rdf #RSD: rsd.xml #Search: cgi-bin/search.cgi It seems that if the pointer in robots.txt contains only the location of an index, then that index will in turn point to another file in many cases, meaning yet another fetch of data. And I don't think it likely that you'd have (or want) a list of available feeds included with the location of (or contents of) the FOAF, all contained in the same file. A format like that above allows the pointer to reference an index when an index is appropriate, and a single location when that is appropriate. Then the contents of each file remain the same and we don't need to come up with a format for the index that is pointed to through the robots.txt file. Only the format(s) on the other end, which in this implementation could be varied, since we specify the format in the pointer. Thoughts? -----Original Message----- From: Chad Everett [mailto:yahoogroups-ipyKpt5L2sG2CP6mJM4Q19i2O/[email protected]] Sent: Thursday, October 16, 2003 3:09 PM To: '[email protected]' Subject: RE: [syndication] site-wide metadata discovery The idea of a hook to a "metadata index file" definitely has appeal. We don't clutter the robots.txt, plus we don't specify a particular file name, plus it could actually be used for sections of domains. Consider: #Site-Index: metadata.xml #Index-Format: http://purl.org/ocs/directory/0.5/ #Allow-Recursion: true Could tell an application that the site index exists in the file named metadata.xml, is in ocs 0.5 format and that the file may exist at each level of the domain (example.com/metadata.xml and example.com/folder/metadata.xml), so if browsing something other than the root, look for the file as it might be there with some different values. Meanwhile: #Site-Index: metadata.xml #Index-Format: http://purl.org/ocs/directory/0.5/ #Allow-Recursion: false Says the same thing, but that there is no recursion possible. Even if metadata.xml exists at lower levels of the domain, don't bother looking for it, as we are telling you not to do so. Some larger sites might like this sort of feature as it could limit requests for data. Another advantage to this is that the index itself could also be provided in different formats, allowing for preferred formats or even selective parsing of the data. If the application sees: #Site-Index: metadata.xml #Index-Format: http://purl.org/ocs/directory/0.5/ #Site-Index: metadata.opml #Index-Format: http://www.opml.org/spec And the application doesn't support opml or ocs, it can just not bother fetching more data because it won't be able to parse the data anyway. End of random reaching for files that might or might not be useful. Of course, there will be some bad behavior where the doc doesn't match the format, or the app grabs the data anyway, but I don't know that that can be prevented, and this ought to help that situation greatly. No excuse for bad behavior if you see something and still choose to get it, even if you can't parse the data within... Naturally this also raises two more questions: - What is the default for recursion? True or false? False may mean less traffic. True may mean more results (also could mean more errors!). - What the heck format will support storing all the metadata for a site? Consider that it might contain feeds, but other useful information could be in there as well - subscriptions, location of foaf, pointer to rsd, etc... -----Original Message----- From: Danny Ayers [mailto:[email protected]] Sent: Thursday, October 16, 2003 2:46 PM To: [email protected] Subject: RE: [syndication] site-wide metadata discovery Nah, I think a little hook is elegant. Your use of Yahoo! Groups is subject to http://docs.yahoo.com/info/terms/