Re: Dynamic vs. Fixed World Views was Re: MARCXML to Topic Maps? MODS to Topic Maps?
Alexander Johannesen <[email protected]>
| Newsgroups | gmane.text.xml.xtm.general |
|---|---|
| Message-ID | <[email protected]> |
marijane white <[email protected]> wrote: > So? This is a problem with lots of data sets. Some folks, say Cory > Doctorow, might say it's a problem with all of them. [1] It seems unfair to > dismiss the utility of MARC for this reason. I've never said it can't be done with MARC. I'm only saying that on a scale of 1 to 10, where 10 is easy, perfect, happy meta data, MARC is a 1 or a 2. It's just harder *because* of the shere complexity of legacy and history. > People still manage to find a > way to make use of their metacrap despite all the very real problems with > it. And, for what it is worth, both Open Library and especially LibraryThing *do* this, and quite well. However, they are not part of the library world, they are external players who might or might not share their findings. > I might argue that librarians have known about these problems longer > than most and that historically they have done a better job than most > because they are at least following principles of description [2] that aim > to limit the impact of said problems, however inconsistently they may be > applied and however bizarrely they choose to do it. Plus, I would think the > inconsistency of the contents are not the fault of MARC but rather the > inconsistent application of the byzantine cataloging rules. In all due respect, there's not *much* difference between MARC and MARC21, and saying that the cataloging rules are somewhat distinct from the MARC wrapper is splitting hairs. But I often rectify this by talking about the Culture of MARC just to be on the safe side. :) I'm fairly confident that one uses MARC outside the AACR2 scope in any capacity. > But anyway. Compare the contents of MARC records as a whole to the contents > any other large scale user-generated data effort. Wikipedia. Musicbrainz, > and the database that inspired it. IMDB. The TV Tropes wiki. Reference > citations. [3] Or commercial datasets -- like Amazon's catalog (quick, go > search Amazon for Mark Twain, then do another search for Samuel Clemens, and > then compare and ponder the results). I suspect you will find identity > issues and examples of all the problems Doctorow cites everywhere. Of course we do. And no one has disputed these things. However, you would expect that the MARC set which was put together by highly trained meta data workers would stand up to this kind of scrutiny. But too often they don't. I guess that's something I'm trying to say as well; the library meta data set is actually not any better[1] than others, even though I expected them to be. [1] Note that "better" is subjective here, and could mean a number of things. It could be better at identity management, which it isn't. It could be better at subject headings, which it might be, but there's a problem there too of ontological matching content to concepts, and arcane rules of how many subject-headings to an item at given times throughout history which impacts semantics heavily. However, the MARC set is substantial, complex, and for some items it is just great! However, you still get lots of false positives, false negatives and outright omitted items that perhaps *should* be direct hits. > As such, > I hope you can forgive me for failing to see why this is a problem for MARC > in particular. Unless you mean to imply that having an inconsistently > applied set of byzantine cataloging rules makes the problem somehow worse. > Does it? I'm puzzled by this. Wouldn't it? Regards, Alex -- Project Wrangler, SOA, Information Alchemist, UX, RESTafarian, Topic Maps --- http://shelter.nu/blog/ ---------------------------------------------- ------------------ http://www.google.com/profiles/alexander.johannesen ---