Re: Slightly OT: MARC

Alexander Johannesen <[email protected]>
Newsgroups gmane.text.xml.xtm.general
Message-ID <[email protected]>
Lars Heuer <[email protected]> wrote:
> Don't tell me that we could merge book X with book Y if they have the
> same author or the same title in Wikipedia....

You're dangerously close to reality ... :)

These are the match / merge processes I've talked about before which
most big organisations apply. Take something like Libraries Australia
which tries to be the union catalog (and don't flame me; it's a rough
comparison) of all Australian libraries. The process is quite complex;
first, upload all MARC records into a database (or new incoming ones,
usually uploaded through FTP from various libraries around the
country), then match records based on a very complex set of rules (a
bit of LOC numbers, the NLA record number, a pinch of ISBN [but only a
little pinch, I think], author fields, titles, serial numbers, local
id's, perhaps try a little FRBR magic (although I don't think they've
reached this stage yet) and so forth, then let it simmer overnight or
so, and dump results into second database, and use that as your
up-to-date database. (And yes, there's more details to this)

>> Given the amount of time and energy that has been devoted to the sub-set of
>> subjects that are books, does that give a clue as to what more ambitious
>> projects to give all subjects a single unique identifier are going to wind
>> up?
>
> Well, one  question remains: Where are the subjects? Where are the
> global identifiers? If there are no global identifiers, all attempts
> to create a topic map from MARC are pretty useless. If neither ISBNs
> nor any other global identifier are given, who could we assume that we
> talk about the same subject?

I think you're starting to see the problem, and why I whine so much
about the lack of identity management in libraries and MARC
specifically. :)

However, to confuse us further, there is something called LCSH
(Library of Congress Subject Headings ;
http://en.wikipedia.org/wiki/Library_of_Congress_Subject_Headings)
which is a controlled vocabulary version of tagging that librarians
use for denoting subjects (more like categories). It's interesting in
that you can create compound identifiers for subjects and create
faceted categories and such, but it is a massive set of subjects
(hundreds of thousands of subjects), and these are the defacto
standard in the library world for this kind of stuff. However, there's
a number of problems with these as well, from cognitive dissonance
between subjects and books (historical reasons, mostly), and
taxonomical (from it's own complexity).

You could derive subjects from this, but not for books. Some times
match / merge routines will use the outer taxonomical level of LCSH
subjects in books as to further match it with other records that might
hold similar subjects, but there are *no* GUUID's in the library
world; the closest you'll get is LOC numbers.

They struggle with this very problem in the FRBR work; if two items
are the same, what is the identifier that combine them? Again I've
lobbied for an infra-structural way to automate this process, but I
feel the generic lack of understanding of the problems described and -
perhaps equally important - the lack of resources and funds is halting
progress in this area.

Funny part is that MARC itself as a wrapper has room for incorporating
new identifiers, so all that needs to happen is for the library world
to recognize the problem and adopt a solution to it.


Regards,

Alex
-- 
 Project Wrangler, SOA, Information Alchemist, UX, RESTafarian, Topic Maps
--- http://shelter.nu/blog/ ----------------------------------------------
------------------ http://www.google.com/profiles/alexander.johannesen ---
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.