Re: Subjects, sets and identifications was: Slightly OT: MARC
Murray Altheim <[email protected]>
| Newsgroups | gmane.text.xml.xtm.general |
|---|---|
| Message-ID | <[email protected]> |
On 20/07/10 23:19, Patrick Durusau wrote:
> Greetings!
>
> I wanted to pull together several items from the discussion. What you
> see below is also going to appear on my blog.
>
> Subjects, Sets and Identifications
[...]
Hi Patrick,
I've kept quiet on this subject despite having a fair bit to say (my
apologies for the rambling nature of this message), as this very much
aligns with that now long-dead project I presented in Montreal at
Balisage 2008. The core of that project was merging the graph
produced by the formal subjects within a library classification
system with the informal graph produced by users of that system,
using Scope to keep them distinct in the merged graph.
Now, we've certainly demonstrated we can bitch till the cows come home
about the state of library classification metadata, but it's really a
pointless debate for this community. If we're talking about libraries
that's what we have available; we're not going to be the ones to clean
up that problem, where the issues are likely more to do with funding
than any lack of will on the part of the library community, though
the twin specters of proprietary software and classification system
licensing cannot be ignored. Yes, everyone should fight to open up
that metadata. Libraries are largely public-funded institutions, so
that information should by rights be owned by the public. That the LoC
and DCMI hold proprietary information is criminal. But even if library
systems are locked down we can often still map them, so long as we can
gain access to even the barest parts of their metadata records, and we
increasingly can.
500 lb cheese
http://find.natlib.govt.nz/primo_library/libweb/action/display.do?ct=display&doc=nlnz_tapuhi923014&indx=1&dum=true&dscnt=0&indx=1&vl%2841331690UI1%29=all_items&srt=rank&tab=default_tab&vl%282087478UI0%29=any&ct=search&frbg=&vid=NLNZ&vl%281UI0%29=contains&fn=search&dstmp=1279631502869&vl%28freeText0%29=500%20lb%20cheese&mode=Basic&scp.scps=&showPnx=true
Ugly. [If that wraps try: http://tinyurl.com/256q83w ]
Whether those records are expressed in MARC, MARC XML (why didn't
they have some imagination and call it MARX?), MODS, RDA, METS or
whatever, that's not a barrier. And as you seem to say, the more
source systems, the more identifiers the merrier: each one provides
a firmer identification of a given entity. If any problems arise we
always have the Scope of the identifier to discern and perhaps even
rank the source. Some sources will be better than others.
And if the identifiers (or should we more correctly say identifier
policies and/or practices) for a given entity are "broken", well, no
matter. It's what the sources themselves are using. If we map an
identifier along with some other commonly-available key terms (e.g.,
title, creator, date) we have *plenty* to canonically identify the
entity even if the identifier we use internal to our Topic Map must
remain local/hidden (i.e., that "binding point" we talked about way
back in Paris).
Identification of the entity's subject can certainly be problematic,
but again, we take what we have available in the (theoretically)
controlled vocabulary of the MARC fields, and taking a Google-like
approach, back-mapping all entities that are co-subject-identified,
again we have plenty to work with.
I think part of my confusion in following this discussion is that it
seems sometimes people want to actually mirror the contents of the MARC
record in the Topic Map rather than just pointing at it (i.e., mapping
it). Or perhaps I'm reading things wrong. But that seems like a fool's
errand. All that's really needed to generate the map is to be able to
canonically identify an entity and then be able to address it. The
less in the map the better; I don't think we're trying to *replace*
the library's system. Given that the library system already has a way
of addressing the entity for at least the short term, the requirements
for mapping via a Topic Map are there laid out on the table for us to
work with. The librarians have left us quite a lot to work with
compared to many domains, cf. zoology or law.
Admittedly, persistent identifier systems are not widespread, so the
identifiers used by the map cannot be relied upon over the long term
and may need to be refreshed, even frequently. One of the projects in
my 'IN' box for this year is a persistent identifier system. It's
been in my 'IN' box for two years now, but we had yet another
reorganisation. I think that's now three in four years, with the one
we're now enduring to last at least another two years. Duh.
[I sometimes yearn for a good solid totalitarian state where we could
take all these idiots out and shoot them.]
If the library system has already been FRBRized, well great. If not,
then the external map may be able to provide the basis for that. In
fact, given that the library records themselves are going to provide
the basis for FRBRization, it could be a Topic Map based application
that actually performs that trick. Hint: if someone were to develop
an application that could read MARC records and generate a FRBRized
hierarchy, that would be an enormous boon to both the Topic Map
community (in demonstrating the viability of the technology), and
to the library community (in providing a tool that is sorely needed).
If I ever get the time to (re)start the Neocortext project, that's
exactly what it's meant to do. Unfortunately my day job is hardly
supportive of that work, as there's little vision left in the
organisation; they're too busy reorganising their reorganisations.
FRBR is still relatively new in the library world (12 years is not
a long time for libraries) and is still being rolled out, and not
always very well. For example (not to pick on a proprietary vendor,
but with expensive software one might expect it to work properly),
Ex Libris' Primo provides support for "FRBRizing" ingested records
by storing the record's FRBR relations as facets in the search
index rather than actually representing a true FRBR structure,
since the database itself is entirely flat. Kinda lame, really,
and it's really only FRBR in name. Check out the <frbr> section in
the URL mentioned above if evidence is needed.
But if we stop expecting library systems to provide that structure
and use an external application (web service) to provide it as part
of an external map of that territory, well, we actually might look
upon that as a Good Thing, at least until library systems can then
take the generated map and incorporate the expressed FRBR
relationships for a given entity into its record. But I'm not even
sure that's a good idea. I kinda like the idea of not forcing upon
the entity's record the requirement that it know its place in the
FRBR hierarchy. The only part of the FRBR tree that need be kept
in the record is the information pertinent to the Item itself,
pointing out into the external map for the Manifestation, Expression
and Work. Having all that in the library database means duplication
of information and no sense of an "object orientation" in the
design.
I.e., unless the library database can itself express the FRBR
hierarchy. I think that's asking too much of library systems,
at least for the next decade or two. Libraries are strapped for
cash, and just today Amazon announced they'd sold more eBooks
than hardcovers. I'd probably say "so what?" given hardcovers
are only a tiny fraction of the softcover market, but point is,
here's a ripe avenue for Topic Maps, and we have a significant
number of people in our community with either a library
background or who actually work in libraries, so let's stop
bickering about the broken library metadata that we can't fix
and actually provide a solution for them. I thought part of
the point of Topic Maps was to be able to help people navigate
the muck, not just complain about it...
Murray
PS. for some reason I receive none of Alex' posts, only peoples'
replies to him. I have no idea why.
...........................................................................
Murray Altheim <murray10 at altheim dot com> === = =
http://www.altheim.com/murray/ = = ===
SGML Grease Monkey, Banjo Player, Wantanabe Zen Monk = = = =
Boundless wind and moon - the eye within eyes,
Inexhaustible heaven and earth - the light beyond light,
The willow dark, the flower bright - ten thousand houses,
Knock at any door - there's one who will respond.
-- The Blue Cliff Record