Re: Slightly OT: MARC
Alexander Johannesen <[email protected]>
| Newsgroups | gmane.text.xml.xtm.general |
|---|---|
| Message-ID | <[email protected]> |
Lars Heuer <[email protected]> wrote: > Don't tell me that we could merge book X with book Y if they have the > same author or the same title in Wikipedia.... You're dangerously close to reality ... :) These are the match / merge processes I've talked about before which most big organisations apply. Take something like Libraries Australia which tries to be the union catalog (and don't flame me; it's a rough comparison) of all Australian libraries. The process is quite complex; first, upload all MARC records into a database (or new incoming ones, usually uploaded through FTP from various libraries around the country), then match records based on a very complex set of rules (a bit of LOC numbers, the NLA record number, a pinch of ISBN [but only a little pinch, I think], author fields, titles, serial numbers, local id's, perhaps try a little FRBR magic (although I don't think they've reached this stage yet) and so forth, then let it simmer overnight or so, and dump results into second database, and use that as your up-to-date database. (And yes, there's more details to this) >> Given the amount of time and energy that has been devoted to the sub-set of >> subjects that are books, does that give a clue as to what more ambitious >> projects to give all subjects a single unique identifier are going to wind >> up? > > Well, one question remains: Where are the subjects? Where are the > global identifiers? If there are no global identifiers, all attempts > to create a topic map from MARC are pretty useless. If neither ISBNs > nor any other global identifier are given, who could we assume that we > talk about the same subject? I think you're starting to see the problem, and why I whine so much about the lack of identity management in libraries and MARC specifically. :) However, to confuse us further, there is something called LCSH (Library of Congress Subject Headings ; http://en.wikipedia.org/wiki/Library_of_Congress_Subject_Headings) which is a controlled vocabulary version of tagging that librarians use for denoting subjects (more like categories). It's interesting in that you can create compound identifiers for subjects and create faceted categories and such, but it is a massive set of subjects (hundreds of thousands of subjects), and these are the defacto standard in the library world for this kind of stuff. However, there's a number of problems with these as well, from cognitive dissonance between subjects and books (historical reasons, mostly), and taxonomical (from it's own complexity). You could derive subjects from this, but not for books. Some times match / merge routines will use the outer taxonomical level of LCSH subjects in books as to further match it with other records that might hold similar subjects, but there are *no* GUUID's in the library world; the closest you'll get is LOC numbers. They struggle with this very problem in the FRBR work; if two items are the same, what is the identifier that combine them? Again I've lobbied for an infra-structural way to automate this process, but I feel the generic lack of understanding of the problems described and - perhaps equally important - the lack of resources and funds is halting progress in this area. Funny part is that MARC itself as a wrapper has room for incorporating new identifiers, so all that needs to happen is for the library world to recognize the problem and adopt a solution to it. Regards, Alex -- Project Wrangler, SOA, Information Alchemist, UX, RESTafarian, Topic Maps --- http://shelter.nu/blog/ ---------------------------------------------- ------------------ http://www.google.com/profiles/alexander.johannesen ---