Re: Dynamic vs. Fixed World Views was Re: MARCXML to Topic Maps? MODS to Topic Maps?

Alexander Johannesen <[email protected]>
Newsgroups gmane.text.xml.xtm.general
Message-ID <[email protected]>
Hiya,

Patrick Durusau <patrick-Q/[email protected]> wrote:
> As far as librarians not making their "...non-normalized untype prose
> understandable to the rest of us..." I suspect that many librarians could
> say the same things about CS literature.

The difference here is that the library world is a niche that aren't
cut off from the CS literature, and "CS literature" isn't exclusive to
anyone. The divide you describe is artificial and self-afflicted
one-sided.

> I don't think complaining about what other people are failing to make
> transparent to us is a useful tack.

What, shut up because the criticism is negative? Surely not?

> Maybe building tools to help others make such fields transparent might.

Making tools takes time and money, and I doubt very much that outside
time and money will be worth the money that can be made in the library
world on meta data. As an academic venture it's still a sexy problem,
but not commercial.

>> I'd love to see the
>> library world actually take this problem seriously, but it seems most
>> librarians are in denial of what impact this little problem have on
>> their relevancy to society.
>
> Why is that a familiar refrain?

Because it's true. :)

> I have heard the claims that community X should spend its time and resources
> on Y because group Z thinks it will lead to democracy, innovation and save
> the planet, all by changing human nature. Yeah, right.

Hang on; remember that I've fought this battle from the *inside*; I
was a librarian once. And I can give you a long list of smart
librarians who say the *exact* same thing. There is a painfully slow
awareness in the library world of these issues, and an even slower
adaption and reaction to it, but I suspect it will be too slow to make
any difference.

> So what is the difficulty in making a topic map from what we do understand
> or can map and allowing others to contribute what they will?

Some fields are simpler than others, but go to your catalog and list
all records with a given field and subfield that *should* be typed (or
at least hit some degree of enumerated value list), and see if it
happens (and my experience here is that unless the field is completely
free-text, this never happens. Never.). Here you've got a choice; only
use the ones that muster *inside* the scope of sane data, or find some
way to fix the problems with those who fall without. And now you
repeat this process for every single field and subfield. This is all
hard enough if there was a defined list of enumerations / values /
schemas, but most of the time there isn't (RDA tries to fix a little
bit of this). It's all fine and well that you've got values in your
fields, but those values needs to have some semantic values, yes?
That's the whole point, is it not?

Even simple things can be excruciatingly hard. Pull out any MARC
record and give me the rules for how to determine whether the item in
question is an eBook. Or a paperback. Or a pamphlet. Various types
have very varying set of rules and fields to match on to get this
remotely right. One of the main things about TM is goodness that comes
from being typeified, so determining type is both important and very
hard. Somewhere between these two there are opportunities to be made,
but again it's a question of return on investment. The *world* don't
get much back from this investment, at least not in their eyes, but
the library world have everything to gain from it. So who should foot
the bill?

> Or to put it more bluntly, why is this all or nothing?

Either you try to do it right, or you're wasting money.

> Part of being "dynamic" is that the knowledge a topic map represents can be
> refined over time.

Just like what librarians have done with the MARC data set over the
last 40 years.

>> Yet the library world is mostly void of understanding Topic
>> Maps, little less implementation of it. I have many online friends in
>> the library world, and they are all as frustrated as me with the lack
>> of a MARC cleanup job that might enable this "many access points"
>> dream. If there *was* such interest I know a handful of very smart
>> people who would jump on it straight away! Alas, there is a distinct
>> disjoint between what the library needs to do and the management that
>> runs it.
>
> You mean the lack of someone else to clean up the data to our liking.

Why is this about "someone else"? I'm talking about cleaning up
library meta data for the benefit of librarians. The world don't
really care all that much about this, and certainly don't understand
the issues involved, probably won't cry if the library world
disappears alltogether (to them, it's the knowledge which is
important, not the box it comes in or the place they gain it). The
library needs to take some responsibility for their own meta data,
make it ready for the world, otherwise the world will simply go on
making its own and ignoring what the library has spent 200 years
building up. I would think this was terribly clear cut.

> But, what if everyone cleaned up the part of greatest interest to them? I
> spend a lot of time with older CS literature so I might try my hand at some
> of those records. Other people have other areas of interest.

I have myself cleaned up tons of MARC regarding early music (and
specifically the context around Claudio Monteverdi), but it's a futile
exercise in the long run because new meta data suffers from the
previous errors. If you're planning on continuous integration of MARC
meta data you need a framework for filtering, matching, killing,
fixing and handling the meta data. The stupid thing is that any major
library institution, private or public, in the world!!! has got the
equivalent of this already in place, OCLC, LOC, national libraries
around the world (and I've got special knowledge of the Libraries
Australia), all huge filtering systems costing millions in people and
resources. Why aren't these efforts simply open sourced? Why aren't
these projects openly discussed, shared and extended?

> Rather than seeing the problem as an all or nothing Mount Everest of data
> conversion, all I am suggesting is that as people use a topic map based on
> such data, some of them would have enough interest to improve the topic map
> by contributing to it. Whether than would be enough or not, I honestly don't
> know.

Any collection of MARC you fix up will be doomed to live outside the
original context for its useful lifetime. I think you gravely
underestimate the problem of MARC, but I appreciate your positive
enthusiasm. :) I was once there, too.

> I do know that waiting for a library messiah to come along and convert
> decades worth of data to some unspecified level of quality is even less
> likely to be effective.

I've given up on the library world. They will not be able to save
themselves from going under, and the library will slowly turn into an
archive of objects of peripheral interest. But that's just little
positive me talking.


Regards,

Alex
-- 
 Project Wrangler, SOA, Information Alchemist, UX, RESTafarian, Topic Maps
--- http://shelter.nu/blog/ ----------------------------------------------
------------------ http://www.google.com/profiles/alexander.johannesen ---
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.