Re: Dynamic vs. Fixed World Views was Re: MARCXML to Topic Maps? MODS to Topic Maps?

Patrick Durusau <patrick-Q/[email protected]>
Newsgroups gmane.text.xml.xtm.general
Message-ID <[email protected]>
Alexander,

On 7/18/2010 1:50 AM, Alexander Johannesen wrote:
> Aki and Patrick;
>
> Ok, this just might turn a bit ranty. Stand clear.
>
>    
My, my, having read your post you are in an ranty mood. ;-)

Realize I have been laboring for over twenty years now to get biblical 
scholars to use XML to that at least their scholarship can be preserved. 
Haven't reached the issues of greater re-use, collaboration, etc. That 
is two or three life times away. ;-) If a tenth of one percent are using 
XML I would be surprised.

So, I am no stranger to long efforts in the face of adversity. But 
anything worth doing isn't going to be easy. ;-)

> I'm increasingly frustrated in this conversation by the lack of
> real-world applicability, and I don't say this lightly nor do I mean
> any disrespect, but the MARC dataset is crap. There, I said it. It's
> rubbish for any other purpose than to be worked with and searched by
> trained librarians. I'm sure Patrick will jump in and tell me off for
> being so dismissive, that I shouldn't hinder innovation and progress
> by being so rooted in reality. Well, tough. Deal with it. Deal with
> the fact that the library meta data set is rubbish, and if you want to
> search for what golden nuggets there might be inside, feel free to do
> so only expect it to be hard. Real hard. Not just technically hard,
> but semantically, ontologically and identifiably hard. Like crazy
> hard. But if you're happy with crazy hard, I'm not going to encourage
> anyone to stop doing so, by all means.
>
>    
You really should ask before presuming what I will or will not say/agree 
to, etc.

See my posting, Subject Headings and the Semantic Web, 
http://tm.durusau.net/?p=14, where I point out that Library of Congress 
subject headings aren't correctly used by *44% of librarians.* BTW, the 
researcher in that case, Karen Marley, has been working on this topic 
for more than 20 years. Has the research to prove she is correct. 
Repeatedly.

You can visit the Library of Congress website and search their catalogs, 
www.loc.gov and judge for yourself how much her research has been heeded.

But, she hasn't thrown in the towel and neither will I.

I freely grant that everything Alex says is true. I find being rooted in 
"reality" as Alex puts it rather comforting.

But my "reality" does not include Alex's despair. Perhaps that is the 
difference.


> It's not that I'm trying to be difficult, but I *have* spent years of
> my life working in this exact problem space, converting MARC into
> anything else that might be useful, and especially into Topic Maps.
> Maybe I'm doing it wrong, maybe I'm not really all that good at it,
> maybe I'm just not understanding something basic, and feel free to
> prove me wrong, I would be delighted if you did! In fact, I dare you!
>
>    

I think the single person, group, etc. approach is wrong. It is too 
top-down to succeed. Mostly because no one person or group has the 
expertise to make the mapping work.

See my posting about Alan Bosworth's article in that regard: 
http://tm.durusau.net/?p=1126

The single author approach to topic maps is another failure of the 
community.

How can we represent the semantic diversity in any community by having 
one group or person author a topic map?

Answer: We can't.

Objection: Oh, but then it won't be consistent with its design, etc.

Answer: So? People are inconsistent and I would much rather find the 
places where we are inconsistent or use different terminology as the 
basis for further conversation and clarification. May be and remain 
unresolved.

The problem is that for all of its power to include diversity, we have 
reverted to silo like thinking in the construction of topic maps.

Witness a well know identity server that only allows the domain where an 
identifier originates to say what are equivalent identifiers. That's 
just bullshit. How could the originator of an identifier know all the 
alternative identifiers for the subject their identifier identifies? And 
what if they object?

> Aki ;
>    
>>> topic maps back to MARCXML without loosing information. Alex and all you
>>> other library people, would you consider this kind conversion useful?
>>>        
> No, not at all. MARCXML is evil, and should be put out of its misery.
> It was a first go at being cool and open with the world, a first
> miserably poor go. I've written about before, too ;
>
>     http://shelterit.blogspot.com/2008/09/marcxml-beast-of-burden.html
>
> Notice the comments by trained librarians at the end as well. Or do a
> search around the intertubes for how much love MARCXML should get.
> Sure you can get a roundtrip of meta data through doing this, but why?
> What's the purpose of roundtripping rubbish meta data? MARCXML does
> not do anything different than MARC. There's nothing to gain. Even the
> familiarity it gives is a kludge on the job that needs to be done.
>
> I have for many years endured library systems that can't even get
> their simplest of identity management ideas right, the idea of using
> LOC numbers (and try finding a true definition of what that means :)
> as a point to match and merge stuff. Laughable. Same with authority
> records which are all parsing and text-based character comparisons.
> Projects like People Australia and OCLC's Identities are meant to be
> some kind of answer to this, yet neither is a) supported in MARC, and
> b) still gets it wrong more often than right. The *only* thing done
> right internally is the focus on FRBR, a standard and open model for
> all things in the bibliographic world (but don't get me started on the
> FRBR use-cases that stand at the centre of the effort, and how
> outdated they are! FRBR was released in 1993 in Stockholm, and is
> *still* only on the prototype level)
>
> Look, I understand what Patrick is saying, I'm not *that* stupid. :)
> And all this data is certainly ripe with golden nuggets. They're
> there, I know it, I've seen them. But to get to them you need to
> extract them, and *that* is where the problem lies. You need to make
> sure that one piece of prose is the same as some other piece of prose,
> and there is *no* easy way to do that. In fact, the only people who
> have the expertize to pull that off are the librarians. This is why I
> say librarians should clean up their own data, because anyone else is
> on a sure path to fail (who else can appreciate the many layers of the
> concept "title" they got. In the real world we got "title" and a
> possible "subtitle", and that's it. Not librarians; they've got
> hundreds of "title" things to worry about. Here, knock yourself out;
> http://www.loc.gov/marc/bibliographic/bd20x24x.html How are we
> supposed to create any Topic Maps to MARC conversion with such a huge
> and highly specialized model at hand? Sure, it's doable, but not by
> mere mortals). this is not rocket-science to see, nor should it be
> controversial, nor should there be any grounds for disagreement with
> such. Librarians are experts in library meta data. Anybody else will
> struggle more than a librarian fixing it up. This is just, you know,
> life. It's how the world works. But if the librarians *don't* clean it
> up, who will? Who's got the knowledge and expertize? And at what
> price? What is the *value* of cleaned up bibliographic data? Can it be
> measured in a way that leverages it for more than very specific
> sub-sets of researchers that find these things fascinating? Does the
> world want bibliographic meta data as much as they want information
> and / or knowledge and / or wisdom?
>
> I don't understand the disagreement from a purely pragmatic point of view!
>
>    
I don't disagree that MARCXML could be improved, but what's your point?

I can't remember ever seeing an XML format that I didn't think could be 
improved. ;-)

Well, some more than others and clearly MARCXML was a first attempt that 
needs a lot of work.

We could just as easily, in all likelihood, work from the MARC record 
format.

I think you are way too hung up on the presentation of the data. Sure, 
MARC records are difficult to read, rely on whitespace, etc.

All of that is presentation. Don't like that presentation, write another 
one and view it that way.

Where I strongly disagree is with the statement:

> You need to make
> sure that one piece of prose is the same as some other piece of prose,
> and there is *no* easy way to do that. In fact, the only people who
> have the expertize to pull that off are the librarians.

Sorry, that is the sort of provincial bullshit that created the problem 
in the first place. Lawyers, biblical scholars, network/Unix 
administrators (to name three groups I have been a member of long enough 
to qualify as an insider) all make the same claim. As do all other 
professions. And non-professions such as government agencies, research 
projects, oil companies, taxi drivers, etc.

And they are all correct, to a degree. But, waiting for them to initiate 
mappings is clearly not working.

Do we agree on that much?

I will be saying more about this in a series of blog posts I am 
composing now, but we need to talk about the *offensive* use of topic maps.

You know the revised saying:

If you can see it, you can identify it.

If you can identify it, you can map it.

If you can map it, you can hit it.

If you can hit it, ....

Well, that's the one I want to use to sell topic maps to military and 
intelligence agencies.

For civilian operations just change the third line to "...you can merge 
it." And leave the fourth line out.

Agreement by the target is not required.

What I envision is wholesale conversion (or better) of data into topic 
maps with subject recognition and crowd sourcing the correction.

Sure, lots of it will be wrong so there needs to be an easy correction 
mechanism for *all the disgruntled* librarians you know of to do 
something more effective that bitching to each other about management 
not changing.

So management isn't going to change? So what? You do know the definition 
of insanity? Doing the same thing over and over and expecting different 
results. Apply that to bitching about management and see what you think.

>> As you point out, it gives librarians something familiar, which is always a
>> good thing when trying to interest someone in an "improvement" of any kind.
>>      
> What good would this possibly do? What's the point of saying that
> their MARC can now also be found, as is, but in a different container?
> I don't get it. They got it in various containers, including RDF,
> CoINs, MODS/MADS (admittedly with model modifications) and the
> 'lovely' MARCXML. The improvements of Topic Maps are the stuff that
> you don't find in the old meta data at all, like typed data and
> identity management. And this gets us to the point I was raising
> before; either you pull the meta data out, and clean it up and make it
> pretty and deal with the fact that this is now a new data set, all
> nice and clean, that bares no resemblance to its original and can no
> longer be associated with it, or, you need to make sure that the
> back-end flow of meta data now deals with the glorious new changes you
> propose. Either solution is, to put it mildly, not very good or
> plausible.
>
> As an example, have a look at this ;
>
>     http://nationaltreasures.nla.gov.au/
>
> This is a Topic Maps website where the initial process was the following ;
>
>     1. Do MARC search and extract MARC records of the items that were
> to go on display
>     2. Give MARC records to developers to create glorious website
>
> Once we got the meta data we tried to do a rather clean conversion. It
> failed miserably. No, the public didn't complain, the librarians
> complained, because the clean conversion meant that the meta data
> display of the items were all, um, screwed up, bibliographically
> speaking, they didn't take into account some of the peculiar
> attributes of, say, the Title Statement, nor did they cite in the
> right format, nor did they portray authorship correctly. So we needed
> to clean it up properly. And then we made it work. And *then* they
> gave us a batch of MARC records as an update to the site, and we had
> to do the whole thing over again (and continuous integration issues
> ensued!). This dance go back and forth, between the base data and its
> ontology / model (MARC) and the Topic Maps cleaned and typeified
> version. The cleaner it gets, the harder it is to integrate, the more
> complex the cleanup process. And then you've got all the points of
> reference that you need to point that cleanup process to. WikiPedia?
> DBpedia? Any of the other hundreds of RDF ontologies? Or, how about
> the FRBR ontology, with added FRBR?
>
> So in short, MARCXML to Topic Maps might give them something familiar,
> and it still will require miracles. And that's what I'm questioning;
> the value of that familiarity. I've done that before, and it doesn't
> always help the process. In fact, often it destroys the process,
> because if you claim one model is somewhat supported by another,
> people lose sight of the fact that they need to completely convert to
> a new model. The familiarity has killed off more than one project.
>
>    
Well, if you screwed up the entries, bibliographically speaking, I can 
imagine that it got off to a rocky start.

Yes, cleaning up data, from a subject identification perspective, is 
always going to be hard work.

When did I ever say otherwise?

It may be that familiarity with MARCXML isn't worth the effort.

Perhaps it should start from MARC itself.

BTW, what is killing your attempts can be summarized with "...they need 
to completely convert to a new model."

Yeah, right. I can see that happening. Do you just like failing?

With topic maps we can represent their "current" model as well as add 
other models, including the "new" one that you think is better.

Topic maps are *not* a rip-n-replace technology. Multiple models can 
exist side by side quite comfortably.

True, to write a topic map that I would find useful for library data it 
would have a lot more explicit subjects than say a MARC record but so what?

If a librarian wants to *not* see the explicit subjects that I added and 
want to view a MARC record with space delimited formatting, that should 
be an option. MARC records are a *presentation* issue.


> ...
>
>    
>> At this point I am sure Alex is going to object that such strings are used
>> inconsistently, etc. Which is very true. But, since our options are to curse
>> the library community for not normalizing decades if not centuries worth of
>> data or identifying and then refining our identifications, I am arguing for
>> the latter.
>>      
> Nonsense, these are not our only options, nor is the former what I
> have suggested we do (even though it's what I'm doing now, after all
> these years). In fact I have made several suggestions, both pragmatic
> and library-specific, but they do require that people with a minimum
> of knowledge to pick it up and do it, otherwise it is - like I've
> state several times - just an academic exercise.
>
>    

Sure, assign other people work to do at their expense. Why didn't I 
think of that?
>> (Noting that the process of normalization is as fraught with the
>> potential for the same inconsistencies as the processes that created the
>> data in the first place.)
>>      
> That may be so, but in order to find out how that hopeless mass of
> hobbled-together pieces can stand up to normalized scrutiny, you have
> to pull them apart and see if they fit elsewhere. And there's the rub;
> where in the world does this highly bibliographic meta data fit if it
> isn't in the highly bibliographic world? You want to set the data free
> to be useful, right? Remember that "free" means different things in
> different contexts. A freed bibliographic dataset could mean
> absolutely nothing in a world that describes "stacks" as "shelves."
>
>    
Actually I have little or no interest in setting data "free," whatever 
connotation you want to associate with being "free."

My interest is in enabling searches across vocabularies that have 
changed over four millennia and too many language and cultures to be 
accurately enumerated. The bibliographic data gathered by the library 
community touches on issues of access to secondary and sometimes primary 
literature.

As you say, obtaining meaningful access to library bibliographic data 
will require pulling it apart (at least from one perspective) but my 
argument is that if there are advantages to doing so, then the library 
community will follow on its own. If there are not, it won't. But we 
won't know unless we try.

Consider the perennial complaints about commercial vendors of library 
catalog software. But enough libraries keep using them to support their 
continuation. And to have several "open source" projects to replace 
them. Some of which are quite good. Others, well, contact me off-list 
for an example of a perfectly horrid search example. And it is 
unfortunately in use by at least one state library system. I would not 
run it to track my personal collection.

I have gone on too long but I have a story that may be worth the extra 
space:

In the Cross and the Switchblade, one of the people recounts a story of 
how to take a bone away from a hungry dog. You can try to simply take 
the bone away (your approach) and you are likely to get bitten. The bone 
is all the dog has. The alternative is to drop a steak down next to the 
dog. He will drop the bone voluntarily and pick up the steak. (the 
approach I am advocating).

So, rather than bitching about the MARC/MARCXML bone, let's thrown down 
a topic map steak and see if the library community will drop the bone.

> And sorry for being polemic, and I'm sorry yet again for the
> negativity; I've spent too many years in the library world trying
> very, very hard to come up with solutions in this problem-space, and
> this is just a result of all those years of fighting the good fight
> against inertia and silo-mentality (something that *only* librarians
> themselves can fix, mind you).
>
>    
True, only librarians can change their own minds, so should we continue 
to try to take their bone away or throw a steak their way?

> On that note, I think I've said my piece, and I'm backing down now.
> Thanks for listening.
>
>    
You're more than welcome.

Hope you are at the start of a great week!

Patrick
> Regards,
>
> Alex
>    

-- 
Patrick Durusau
patrick-Q/[email protected]
Chair, V1 - US TAG to JTC 1/SC 34
Convener, JTC 1/SC 34/WG 3 (Topic Maps)
Editor, OpenDocument Format TC (OASIS), Project Editor ISO/IEC 26300
Co-Editor, ISO/IEC 13250-1, 13250-5 (Topic Maps)

Another Word For It (blog): http://tm.durusau.net
Homepage: http://www.durusau.net
Twitter: patrickDurusau
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.