Re: Dynamic vs. Fixed World Views was Re: MARCXML to Topic Maps? MODS to Topic Maps?

Alexander Johannesen <[email protected]>
Newsgroups gmane.text.xml.xtm.general
Message-ID <[email protected]>
Hi,

marijane white <[email protected]> wrote:
> If an ILS can parse a MARC record, why can't a MARC record be parsed into a
> Topic Map?

It's not that it can't happen, it's just that it's a hard struggle and
doesn't seem to find momentum to actually happen.

> If they really are as rubbish as you claim, how is it possible
> to do anything at all with them?  How is it that library programmers have
> created things like, say, SolrMarc, and the things built upon them like
> VuFind and Blacklight?

All of these projects do pretty much the same thing ; export out of
your ILS / OPAC your MARC records, filter and / or clean the data,
recreate the MARC model to various degrees, and then isolate the new
data set to make use of it. At no point does the cleaned data get back
into the MARC infra-structure because there is no where for them to
go.

The "rubbish" statement picks more at two specific things with the
MARC data more than anything ;

1. Little to no typed data, no enumeration of values
2. Little to no identity management

Then there's the larger problem of the MARC infra-structure (the
library infra-structure) where the tools that deal with MARC have no
validation of values. This is why I'm so irate about MARCXML because
they claim one of the big advantages is validation, but you can only
validate that the wrapper is sane, not that the content is. And the
content is built by hand (mostly) by librarians for over 40 years. The
infra-structure and the tools do not and can not validate that fields
that are supposed to contain dates or date ranges indeed contain dates
or date ranges, cannot validate that a field which contain, say, a
reference to an author indeed points to an author. Here's the problem
a bit more specified ;

Every field (and I include subfields and indicators here) has a name,
and in the MARC21 standard has a host of rules and best-practices.
However, the rules are very often not something that can be
automatically checked. For example, a field that contains date range
might have "19th - 20th", or "ca. 3 bc.", "1910-1940", "50's",
"classical period" in them. It's well and fine to write special
parsing rules to extract roughly what is meant and then represent that
in some ISO format, but remember that there are thousands of fields,
all different in what they represent, and often you'll find that a
record represents an item with a specific compound of fields (so that
if this field and that field says X and Z, then if field A is this,
then it's a paperback, if not it's a hardcopy, and so on) making it
very hard for typification. (I tried to find one of these official
OCLC or LOC documents that lists the process for just determining what
type of thing any given record is, but my Google-foo is weak this
morning. Suffice to say, it's a fun process!)

Basically MARC is hundreds of free-text fields that needs to be parsed
and checked against some back-end facit. As stated earlier, all large
libraries and organisations *already* have their own projects in place
that do this kind of thing (for a variety of reasons, too, like they
want to clean up the MARC, or for outside import / export reasons),
from NLA (with Libraries Australia), OCLC, LOC, other national
libraries, Library Thing, OpenLibrary, and on and on. But we don't
have one we all can use and collaborate on. What should happen is of
course a project that tries to do all this, an open-source project all
of us should join in order to create this huge mapping, be it
ontologically or software or processes and filter specifications, or
whatever, I'm not fussy, anything will do. But it doesn't exist, nor
does it look like it will be created as too many librarians don't see
the problem, and outsiders will look at it and probably not think it
worth their time. I shudder at the thought of all those resources
wasted because of this situation. If we all value the MARC data set so
much, and we all agree that it should be free (to which they all don't
agree) it should be done, methinks, but then a lot of people seem to
disagree with me on this. :)

> Is the situation really so dire?

I think so. As the library world is becoming less and less valuable to
society at large, the need for their MARC data set decrease as outside
bodies are finding other alternatives instead as the library world
can't or won't be more open about this. Knowledge is moving out of
books and into electronic form where non-librarians can do a better
job at sifting through information in search of the nuggets. The
bibliographic world will slowly turn into museums of special interest
objects. But I don't know how long this will take.

Don't get me wrong, there's always going to be libraries, but they
will not be the bastion of knowledge as they have been for centuries,
and it will be very interesting to watch and see what sacrifices and
compromises that will happen. A lot of librarian values will be lost.

Anyway, that's my positive pep-talk for the day.


Regards,

Alex
-- 
 Project Wrangler, SOA, Information Alchemist, UX, RESTafarian, Topic Maps
--- http://shelter.nu/blog/ ----------------------------------------------
------------------ http://www.google.com/profiles/alexander.johannesen ---
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.