Re: Dynamic vs. Fixed World Views was Re: MARCXML to Topic Maps? MODS to Topic Maps?
Patrick Durusau <patrick-Q/[email protected]>
| Newsgroups | gmane.text.xml.xtm.general |
|---|---|
| Message-ID | <[email protected]> |
Alexander, On 7/18/2010 1:50 AM, Alexander Johannesen wrote: > Aki and Patrick; > > Ok, this just might turn a bit ranty. Stand clear. > > My, my, having read your post you are in an ranty mood. ;-) Realize I have been laboring for over twenty years now to get biblical scholars to use XML to that at least their scholarship can be preserved. Haven't reached the issues of greater re-use, collaboration, etc. That is two or three life times away. ;-) If a tenth of one percent are using XML I would be surprised. So, I am no stranger to long efforts in the face of adversity. But anything worth doing isn't going to be easy. ;-) > I'm increasingly frustrated in this conversation by the lack of > real-world applicability, and I don't say this lightly nor do I mean > any disrespect, but the MARC dataset is crap. There, I said it. It's > rubbish for any other purpose than to be worked with and searched by > trained librarians. I'm sure Patrick will jump in and tell me off for > being so dismissive, that I shouldn't hinder innovation and progress > by being so rooted in reality. Well, tough. Deal with it. Deal with > the fact that the library meta data set is rubbish, and if you want to > search for what golden nuggets there might be inside, feel free to do > so only expect it to be hard. Real hard. Not just technically hard, > but semantically, ontologically and identifiably hard. Like crazy > hard. But if you're happy with crazy hard, I'm not going to encourage > anyone to stop doing so, by all means. > > You really should ask before presuming what I will or will not say/agree to, etc. See my posting, Subject Headings and the Semantic Web, http://tm.durusau.net/?p=14, where I point out that Library of Congress subject headings aren't correctly used by *44% of librarians.* BTW, the researcher in that case, Karen Marley, has been working on this topic for more than 20 years. Has the research to prove she is correct. Repeatedly. You can visit the Library of Congress website and search their catalogs, www.loc.gov and judge for yourself how much her research has been heeded. But, she hasn't thrown in the towel and neither will I. I freely grant that everything Alex says is true. I find being rooted in "reality" as Alex puts it rather comforting. But my "reality" does not include Alex's despair. Perhaps that is the difference. > It's not that I'm trying to be difficult, but I *have* spent years of > my life working in this exact problem space, converting MARC into > anything else that might be useful, and especially into Topic Maps. > Maybe I'm doing it wrong, maybe I'm not really all that good at it, > maybe I'm just not understanding something basic, and feel free to > prove me wrong, I would be delighted if you did! In fact, I dare you! > > I think the single person, group, etc. approach is wrong. It is too top-down to succeed. Mostly because no one person or group has the expertise to make the mapping work. See my posting about Alan Bosworth's article in that regard: http://tm.durusau.net/?p=1126 The single author approach to topic maps is another failure of the community. How can we represent the semantic diversity in any community by having one group or person author a topic map? Answer: We can't. Objection: Oh, but then it won't be consistent with its design, etc. Answer: So? People are inconsistent and I would much rather find the places where we are inconsistent or use different terminology as the basis for further conversation and clarification. May be and remain unresolved. The problem is that for all of its power to include diversity, we have reverted to silo like thinking in the construction of topic maps. Witness a well know identity server that only allows the domain where an identifier originates to say what are equivalent identifiers. That's just bullshit. How could the originator of an identifier know all the alternative identifiers for the subject their identifier identifies? And what if they object? > Aki ; > >>> topic maps back to MARCXML without loosing information. Alex and all you >>> other library people, would you consider this kind conversion useful? >>> > No, not at all. MARCXML is evil, and should be put out of its misery. > It was a first go at being cool and open with the world, a first > miserably poor go. I've written about before, too ; > > http://shelterit.blogspot.com/2008/09/marcxml-beast-of-burden.html > > Notice the comments by trained librarians at the end as well. Or do a > search around the intertubes for how much love MARCXML should get. > Sure you can get a roundtrip of meta data through doing this, but why? > What's the purpose of roundtripping rubbish meta data? MARCXML does > not do anything different than MARC. There's nothing to gain. Even the > familiarity it gives is a kludge on the job that needs to be done. > > I have for many years endured library systems that can't even get > their simplest of identity management ideas right, the idea of using > LOC numbers (and try finding a true definition of what that means :) > as a point to match and merge stuff. Laughable. Same with authority > records which are all parsing and text-based character comparisons. > Projects like People Australia and OCLC's Identities are meant to be > some kind of answer to this, yet neither is a) supported in MARC, and > b) still gets it wrong more often than right. The *only* thing done > right internally is the focus on FRBR, a standard and open model for > all things in the bibliographic world (but don't get me started on the > FRBR use-cases that stand at the centre of the effort, and how > outdated they are! FRBR was released in 1993 in Stockholm, and is > *still* only on the prototype level) > > Look, I understand what Patrick is saying, I'm not *that* stupid. :) > And all this data is certainly ripe with golden nuggets. They're > there, I know it, I've seen them. But to get to them you need to > extract them, and *that* is where the problem lies. You need to make > sure that one piece of prose is the same as some other piece of prose, > and there is *no* easy way to do that. In fact, the only people who > have the expertize to pull that off are the librarians. This is why I > say librarians should clean up their own data, because anyone else is > on a sure path to fail (who else can appreciate the many layers of the > concept "title" they got. In the real world we got "title" and a > possible "subtitle", and that's it. Not librarians; they've got > hundreds of "title" things to worry about. Here, knock yourself out; > http://www.loc.gov/marc/bibliographic/bd20x24x.html How are we > supposed to create any Topic Maps to MARC conversion with such a huge > and highly specialized model at hand? Sure, it's doable, but not by > mere mortals). this is not rocket-science to see, nor should it be > controversial, nor should there be any grounds for disagreement with > such. Librarians are experts in library meta data. Anybody else will > struggle more than a librarian fixing it up. This is just, you know, > life. It's how the world works. But if the librarians *don't* clean it > up, who will? Who's got the knowledge and expertize? And at what > price? What is the *value* of cleaned up bibliographic data? Can it be > measured in a way that leverages it for more than very specific > sub-sets of researchers that find these things fascinating? Does the > world want bibliographic meta data as much as they want information > and / or knowledge and / or wisdom? > > I don't understand the disagreement from a purely pragmatic point of view! > > I don't disagree that MARCXML could be improved, but what's your point? I can't remember ever seeing an XML format that I didn't think could be improved. ;-) Well, some more than others and clearly MARCXML was a first attempt that needs a lot of work. We could just as easily, in all likelihood, work from the MARC record format. I think you are way too hung up on the presentation of the data. Sure, MARC records are difficult to read, rely on whitespace, etc. All of that is presentation. Don't like that presentation, write another one and view it that way. Where I strongly disagree is with the statement: > You need to make > sure that one piece of prose is the same as some other piece of prose, > and there is *no* easy way to do that. In fact, the only people who > have the expertize to pull that off are the librarians. Sorry, that is the sort of provincial bullshit that created the problem in the first place. Lawyers, biblical scholars, network/Unix administrators (to name three groups I have been a member of long enough to qualify as an insider) all make the same claim. As do all other professions. And non-professions such as government agencies, research projects, oil companies, taxi drivers, etc. And they are all correct, to a degree. But, waiting for them to initiate mappings is clearly not working. Do we agree on that much? I will be saying more about this in a series of blog posts I am composing now, but we need to talk about the *offensive* use of topic maps. You know the revised saying: If you can see it, you can identify it. If you can identify it, you can map it. If you can map it, you can hit it. If you can hit it, .... Well, that's the one I want to use to sell topic maps to military and intelligence agencies. For civilian operations just change the third line to "...you can merge it." And leave the fourth line out. Agreement by the target is not required. What I envision is wholesale conversion (or better) of data into topic maps with subject recognition and crowd sourcing the correction. Sure, lots of it will be wrong so there needs to be an easy correction mechanism for *all the disgruntled* librarians you know of to do something more effective that bitching to each other about management not changing. So management isn't going to change? So what? You do know the definition of insanity? Doing the same thing over and over and expecting different results. Apply that to bitching about management and see what you think. >> As you point out, it gives librarians something familiar, which is always a >> good thing when trying to interest someone in an "improvement" of any kind. >> > What good would this possibly do? What's the point of saying that > their MARC can now also be found, as is, but in a different container? > I don't get it. They got it in various containers, including RDF, > CoINs, MODS/MADS (admittedly with model modifications) and the > 'lovely' MARCXML. The improvements of Topic Maps are the stuff that > you don't find in the old meta data at all, like typed data and > identity management. And this gets us to the point I was raising > before; either you pull the meta data out, and clean it up and make it > pretty and deal with the fact that this is now a new data set, all > nice and clean, that bares no resemblance to its original and can no > longer be associated with it, or, you need to make sure that the > back-end flow of meta data now deals with the glorious new changes you > propose. Either solution is, to put it mildly, not very good or > plausible. > > As an example, have a look at this ; > > http://nationaltreasures.nla.gov.au/ > > This is a Topic Maps website where the initial process was the following ; > > 1. Do MARC search and extract MARC records of the items that were > to go on display > 2. Give MARC records to developers to create glorious website > > Once we got the meta data we tried to do a rather clean conversion. It > failed miserably. No, the public didn't complain, the librarians > complained, because the clean conversion meant that the meta data > display of the items were all, um, screwed up, bibliographically > speaking, they didn't take into account some of the peculiar > attributes of, say, the Title Statement, nor did they cite in the > right format, nor did they portray authorship correctly. So we needed > to clean it up properly. And then we made it work. And *then* they > gave us a batch of MARC records as an update to the site, and we had > to do the whole thing over again (and continuous integration issues > ensued!). This dance go back and forth, between the base data and its > ontology / model (MARC) and the Topic Maps cleaned and typeified > version. The cleaner it gets, the harder it is to integrate, the more > complex the cleanup process. And then you've got all the points of > reference that you need to point that cleanup process to. WikiPedia? > DBpedia? Any of the other hundreds of RDF ontologies? Or, how about > the FRBR ontology, with added FRBR? > > So in short, MARCXML to Topic Maps might give them something familiar, > and it still will require miracles. And that's what I'm questioning; > the value of that familiarity. I've done that before, and it doesn't > always help the process. In fact, often it destroys the process, > because if you claim one model is somewhat supported by another, > people lose sight of the fact that they need to completely convert to > a new model. The familiarity has killed off more than one project. > > Well, if you screwed up the entries, bibliographically speaking, I can imagine that it got off to a rocky start. Yes, cleaning up data, from a subject identification perspective, is always going to be hard work. When did I ever say otherwise? It may be that familiarity with MARCXML isn't worth the effort. Perhaps it should start from MARC itself. BTW, what is killing your attempts can be summarized with "...they need to completely convert to a new model." Yeah, right. I can see that happening. Do you just like failing? With topic maps we can represent their "current" model as well as add other models, including the "new" one that you think is better. Topic maps are *not* a rip-n-replace technology. Multiple models can exist side by side quite comfortably. True, to write a topic map that I would find useful for library data it would have a lot more explicit subjects than say a MARC record but so what? If a librarian wants to *not* see the explicit subjects that I added and want to view a MARC record with space delimited formatting, that should be an option. MARC records are a *presentation* issue. > ... > > >> At this point I am sure Alex is going to object that such strings are used >> inconsistently, etc. Which is very true. But, since our options are to curse >> the library community for not normalizing decades if not centuries worth of >> data or identifying and then refining our identifications, I am arguing for >> the latter. >> > Nonsense, these are not our only options, nor is the former what I > have suggested we do (even though it's what I'm doing now, after all > these years). In fact I have made several suggestions, both pragmatic > and library-specific, but they do require that people with a minimum > of knowledge to pick it up and do it, otherwise it is - like I've > state several times - just an academic exercise. > > Sure, assign other people work to do at their expense. Why didn't I think of that? >> (Noting that the process of normalization is as fraught with the >> potential for the same inconsistencies as the processes that created the >> data in the first place.) >> > That may be so, but in order to find out how that hopeless mass of > hobbled-together pieces can stand up to normalized scrutiny, you have > to pull them apart and see if they fit elsewhere. And there's the rub; > where in the world does this highly bibliographic meta data fit if it > isn't in the highly bibliographic world? You want to set the data free > to be useful, right? Remember that "free" means different things in > different contexts. A freed bibliographic dataset could mean > absolutely nothing in a world that describes "stacks" as "shelves." > > Actually I have little or no interest in setting data "free," whatever connotation you want to associate with being "free." My interest is in enabling searches across vocabularies that have changed over four millennia and too many language and cultures to be accurately enumerated. The bibliographic data gathered by the library community touches on issues of access to secondary and sometimes primary literature. As you say, obtaining meaningful access to library bibliographic data will require pulling it apart (at least from one perspective) but my argument is that if there are advantages to doing so, then the library community will follow on its own. If there are not, it won't. But we won't know unless we try. Consider the perennial complaints about commercial vendors of library catalog software. But enough libraries keep using them to support their continuation. And to have several "open source" projects to replace them. Some of which are quite good. Others, well, contact me off-list for an example of a perfectly horrid search example. And it is unfortunately in use by at least one state library system. I would not run it to track my personal collection. I have gone on too long but I have a story that may be worth the extra space: In the Cross and the Switchblade, one of the people recounts a story of how to take a bone away from a hungry dog. You can try to simply take the bone away (your approach) and you are likely to get bitten. The bone is all the dog has. The alternative is to drop a steak down next to the dog. He will drop the bone voluntarily and pick up the steak. (the approach I am advocating). So, rather than bitching about the MARC/MARCXML bone, let's thrown down a topic map steak and see if the library community will drop the bone. > And sorry for being polemic, and I'm sorry yet again for the > negativity; I've spent too many years in the library world trying > very, very hard to come up with solutions in this problem-space, and > this is just a result of all those years of fighting the good fight > against inertia and silo-mentality (something that *only* librarians > themselves can fix, mind you). > > True, only librarians can change their own minds, so should we continue to try to take their bone away or throw a steak their way? > On that note, I think I've said my piece, and I'm backing down now. > Thanks for listening. > > You're more than welcome. Hope you are at the start of a great week! Patrick > Regards, > > Alex > -- Patrick Durusau patrick-Q/[email protected] Chair, V1 - US TAG to JTC 1/SC 34 Convener, JTC 1/SC 34/WG 3 (Topic Maps) Editor, OpenDocument Format TC (OASIS), Project Editor ISO/IEC 26300 Co-Editor, ISO/IEC 13250-1, 13250-5 (Topic Maps) Another Word For It (blog): http://tm.durusau.net Homepage: http://www.durusau.net Twitter: patrickDurusau