Re: Dynamic vs. Fixed World Views was Re: MARCXML to Topic Maps? MODS to Topic Maps?

Carlo Moneti <cmoneti-ai6B2lNGiXhWk0Htik3J/[email protected]>
Newsgroups gmane.text.xml.xtm.general
Message-ID <[email protected]>
To all,

Speaking as a non-expert; accepting the value of MARC data; and accepting the difficulty of MARC conversion:

Isn't much of main MARC data (title, author, publisher, pub date, pub location, classification, etc.) already in structured form, for example, in databases driving library catalogs? And are these not our subjects?

What's more important/useful: converting MARC records to a topic map?; or making MARC data easily findable/discoverable in the search context where it is relevant?

If findability and practical access to MARC data is the main interest in a MARC to topic maps converter, perhaps the main focus should be on a LCSH (or other) library classification to topic maps converter, to which, e.g., book topics derived from MARC data (or existing better structured data) are added, and MARC records added as an occurrence (or broken into multiple typed occurrences where possible, over time).

I imagine that even with unstructured MARC data (e.g. as occurrence), a topic map, as described above, driving an application that offers topic navigation, faceted navigation, and full-text search would make it extremely easy to reliably find/discover those MARC 'gulden nuggets' that relate to the search context.

Cheers,
Carlo Moneti


On 2010.07.18 01:50 Alexander Johannesen wrote:
Aki and Patrick;

Ok, this just might turn a bit ranty. Stand clear.

I'm increasingly frustrated in this conversation by the lack of
real-world applicability, and I don't say this lightly nor do I mean
any disrespect, but the MARC dataset is crap. There, I said it. It's
rubbish for any other purpose than to be worked with and searched by
trained librarians. I'm sure Patrick will jump in and tell me off for
being so dismissive, that I shouldn't hinder innovation and progress
by being so rooted in reality. Well, tough. Deal with it. Deal with
the fact that the library meta data set is rubbish, and if you want to
search for what golden nuggets there might be inside, feel free to do
so only expect it to be hard. Real hard. Not just technically hard,
but semantically, ontologically and identifiably hard. Like crazy
hard. But if you're happy with crazy hard, I'm not going to encourage
anyone to stop doing so, by all means.

It's not that I'm trying to be difficult, but I *have* spent years of
my life working in this exact problem space, converting MARC into
anything else that might be useful, and especially into Topic Maps.
Maybe I'm doing it wrong, maybe I'm not really all that good at it,
maybe I'm just not understanding something basic, and feel free to
prove me wrong, I would be delighted if you did! In fact, I dare you!

Aki ;
>> topic maps back to MARCXML without loosing information. Alex and all you
>> other library people, would you consider this kind conversion useful?

No, not at all. MARCXML is evil, and should be put out of its misery.
It was a first go at being cool and open with the world, a first
miserably poor go. I've written about before, too ;

   http://shelterit.blogspot.com/2008/09/marcxml-beast-of-burden.html

Notice the comments by trained librarians at the end as well. Or do a
search around the intertubes for how much love MARCXML should get.
Sure you can get a roundtrip of meta data through doing this, but why?
What's the purpose of roundtripping rubbish meta data? MARCXML does
not do anything different than MARC. There's nothing to gain. Even the
familiarity it gives is a kludge on the job that needs to be done.

I have for many years endured library systems that can't even get
their simplest of identity management ideas right, the idea of using
LOC numbers (and try finding a true definition of what that means :)
as a point to match and merge stuff. Laughable. Same with authority
records which are all parsing and text-based character comparisons.
Projects like People Australia and OCLC's Identities are meant to be
some kind of answer to this, yet neither is a) supported in MARC, and
b) still gets it wrong more often than right. The *only* thing done
right internally is the focus on FRBR, a standard and open model for
all things in the bibliographic world (but don't get me started on the
FRBR use-cases that stand at the centre of the effort, and how
outdated they are! FRBR was released in 1993 in Stockholm, and is
*still* only on the prototype level)

Look, I understand what Patrick is saying, I'm not *that* stupid. :)
And all this data is certainly ripe with golden nuggets. They're
there, I know it, I've seen them. But to get to them you need to
extract them, and *that* is where the problem lies. You need to make
sure that one piece of prose is the same as some other piece of prose,
and there is *no* easy way to do that. In fact, the only people who
have the expertize to pull that off are the librarians. This is why I
say librarians should clean up their own data, because anyone else is
on a sure path to fail (who else can appreciate the many layers of the
concept "title" they got. In the real world we got "title" and a
possible "subtitle", and that's it. Not librarians; they've got
hundreds of "title" things to worry about. Here, knock yourself out;
http://www.loc.gov/marc/bibliographic/bd20x24x.html How are we
supposed to create any Topic Maps to MARC conversion with such a huge
and highly specialized model at hand? Sure, it's doable, but not by
mere mortals). this is not rocket-science to see, nor should it be
controversial, nor should there be any grounds for disagreement with
such. Librarians are experts in library meta data. Anybody else will
struggle more than a librarian fixing it up. This is just, you know,
life. It's how the world works. But if the librarians *don't* clean it
up, who will? Who's got the knowledge and expertize? And at what
price? What is the *value* of cleaned up bibliographic data? Can it be
measured in a way that leverages it for more than very specific
sub-sets of researchers that find these things fascinating? Does the
world want bibliographic meta data as much as they want information
and / or knowledge and / or wisdom?

I don't understand the disagreement from a purely pragmatic point of view!

> As you point out, it gives librarians something familiar, which is always a
> good thing when trying to interest someone in an "improvement" of any kind.

What good would this possibly do? What's the point of saying that
their MARC can now also be found, as is, but in a different container?
I don't get it. They got it in various containers, including RDF,
CoINs, MODS/MADS (admittedly with model modifications) and the
'lovely' MARCXML. The improvements of Topic Maps are the stuff that
you don't find in the old meta data at all, like typed data and
identity management. And this gets us to the point I was raising
before; either you pull the meta data out, and clean it up and make it
pretty and deal with the fact that this is now a new data set, all
nice and clean, that bares no resemblance to its original and can no
longer be associated with it, or, you need to make sure that the
back-end flow of meta data now deals with the glorious new changes you
propose. Either solution is, to put it mildly, not very good or
plausible.

As an example, have a look at this ;

   http://nationaltreasures.nla.gov.au/

This is a Topic Maps website where the initial process was the following ;

   1. Do MARC search and extract MARC records of the items that were
to go on display
   2. Give MARC records to developers to create glorious website

Once we got the meta data we tried to do a rather clean conversion. It
failed miserably. No, the public didn't complain, the librarians
complained, because the clean conversion meant that the meta data
display of the items were all, um, screwed up, bibliographically
speaking, they didn't take into account some of the peculiar
attributes of, say, the Title Statement, nor did they cite in the
right format, nor did they portray authorship correctly. So we needed
to clean it up properly. And then we made it work. And *then* they
gave us a batch of MARC records as an update to the site, and we had
to do the whole thing over again (and continuous integration issues
ensued!). This dance go back and forth, between the base data and its
ontology / model (MARC) and the Topic Maps cleaned and typeified
version. The cleaner it gets, the harder it is to integrate, the more
complex the cleanup process. And then you've got all the points of
reference that you need to point that cleanup process to. WikiPedia?
DBpedia? Any of the other hundreds of RDF ontologies? Or, how about
the FRBR ontology, with added FRBR?

So in short, MARCXML to Topic Maps might give them something familiar,
and it still will require miracles. And that's what I'm questioning;
the value of that familiarity. I've done that before, and it doesn't
always help the process. In fact, often it destroys the process,
because if you claim one model is somewhat supported by another,
people lose sight of the fact that they need to completely convert to
a new model. The familiarity has killed off more than one project.

...

> At this point I am sure Alex is going to object that such strings are used
> inconsistently, etc. Which is very true. But, since our options are to curse
> the library community for not normalizing decades if not centuries worth of
> data or identifying and then refining our identifications, I am arguing for
> the latter.

Nonsense, these are not our only options, nor is the former what I
have suggested we do (even though it's what I'm doing now, after all
these years). In fact I have made several suggestions, both pragmatic
and library-specific, but they do require that people with a minimum
of knowledge to pick it up and do it, otherwise it is - like I've
state several times - just an academic exercise.

> (Noting that the process of normalization is as fraught with the
> potential for the same inconsistencies as the processes that created the
> data in the first place.)

That may be so, but in order to find out how that hopeless mass of
hobbled-together pieces can stand up to normalized scrutiny, you have
to pull them apart and see if they fit elsewhere. And there's the rub;
where in the world does this highly bibliographic meta data fit if it
isn't in the highly bibliographic world? You want to set the data free
to be useful, right? Remember that "free" means different things in
different contexts. A freed bibliographic dataset could mean
absolutely nothing in a world that describes "stacks" as "shelves."

And sorry for being polemic, and I'm sorry yet again for the
negativity; I've spent too many years in the library world trying
very, very hard to come up with solutions in this problem-space, and
this is just a result of all those years of fighting the good fight
against inertia and silo-mentality (something that *only* librarians
themselves can fix, mind you).

On that note, I think I've said my piece, and I'm backing down now.
Thanks for listening.


Regards,

Alex
-- 
 Project Wrangler, SOA, Information Alchemist, UX, RESTafarian, Topic Maps
--- http://shelter.nu/blog/ ----------------------------------------------
------------------ http://www.google.com/profiles/alexander.johannesen ---
_______________________________________________
topicmapmail mailing list
topicmapmail-Zo64W7twoUFWk0Htik3J/[email protected]
http://www.infoloom.com/mailman/listinfo/topicmapmail
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.