Re: Dynamic vs. Fixed World Views was Re: MARCXML to Topic Maps? MODS to Topic Maps?
Patrick Durusau <patrick-Q/[email protected]>
| Newsgroups | gmane.text.xml.xtm.general |
|---|---|
| Message-ID | <[email protected]> |
Alex, On 7/15/2010 9:48 PM, Alexander Johannesen wrote: > Hiya, > > Patrick Durusau<patrick-Q/[email protected]> wrote: > >> As far as librarians not making their "...non-normalized untype prose >> understandable to the rest of us..." I suspect that many librarians could >> say the same things about CS literature. >> > The difference here is that the library world is a niche that aren't > cut off from the CS literature, and "CS literature" isn't exclusive to > anyone. The divide you describe is artificial and self-afflicted > one-sided. > > On the contrary, many of the verbal differences in CS are entirely "artificial and self-inflicted one-sided." At last count I had 20+ different names for record linkage, most of those strands citing the *same* original 2 or 3 articles that are considered the modern starting point for that technology. Hard to escape from the notion that some of the choices of different terminology were with the intent to develop mini-cults. >> I don't think complaining about what other people are failing to make >> transparent to us is a useful tack. >> > What, shut up because the criticism is negative? Surely not? > > No, not at all. I like to complain. ;-) My concern is that urging other people to do something we lack the will to do ourselves isn't going to produce any result. >> Maybe building tools to help others make such fields transparent might. >> > Making tools takes time and money, and I doubt very much that outside > time and money will be worth the money that can be made in the library > world on meta data. As an academic venture it's still a sexy problem, > but not commercial. > > Only if you fail to realize that effective use of metadata would help the business world avoid the migrations to the next "real" solution to their data management problems. >>> I'd love to see the >>> library world actually take this problem seriously, but it seems most >>> librarians are in denial of what impact this little problem have on >>> their relevancy to society. >>> >> Why is that a familiar refrain? >> > Because it's true. :) > > >> I have heard the claims that community X should spend its time and resources >> on Y because group Z thinks it will lead to democracy, innovation and save >> the planet, all by changing human nature. Yeah, right. >> > Hang on; remember that I've fought this battle from the *inside*; I > was a librarian once. And I can give you a long list of smart > librarians who say the *exact* same thing. There is a painfully slow > awareness in the library world of these issues, and an even slower > adaption and reaction to it, but I suspect it will be too slow to make > any difference. > > Oh, I don't doubt that the profession will change in many ways but the notion that somehow we will all be able to search for information and get useful results, in a world where the amount of data is increasing rapidly, is just absurd. >> So what is the difficulty in making a topic map from what we do understand >> or can map and allowing others to contribute what they will? >> > Some fields are simpler than others, but go to your catalog and list > all records with a given field and subfield that *should* be typed (or > at least hit some degree of enumerated value list), and see if it > happens (and my experience here is that unless the field is completely > free-text, this never happens. Never.). Here you've got a choice; only > use the ones that muster *inside* the scope of sane data, or find some > way to fix the problems with those who fall without. And now you > repeat this process for every single field and subfield. This is all > hard enough if there was a defined list of enumerations / values / > schemas, but most of the time there isn't (RDA tries to fix a little > bit of this). It's all fine and well that you've got values in your > fields, but those values needs to have some semantic values, yes? > That's the whole point, is it not? > > Even simple things can be excruciatingly hard. Pull out any MARC > record and give me the rules for how to determine whether the item in > question is an eBook. Or a paperback. Or a pamphlet. Various types > have very varying set of rules and fields to match on to get this > remotely right. One of the main things about TM is goodness that comes > from being typeified, so determining type is both important and very > hard. Somewhere between these two there are opportunities to be made, > but again it's a question of return on investment. The *world* don't > get much back from this investment, at least not in their eyes, but > the library world have everything to gain from it. So who should foot > the bill? > > >> Or to put it more bluntly, why is this all or nothing? >> > Either you try to do it right, or you're wasting money. > > Ah, and there is the root of the issue. *Right in whose view?* This is what I find disappointing about identifier services that only allow the domain that originated the identifier to say that another identifier is for the same subject. Despite the fact that we could use topic maps to open up library catalogs with their many fields, or allow mappings of arbitrary identifiers, we don't. We are *rebuilding* silos one brick at a time. >> Part of being "dynamic" is that the knowledge a topic map represents can be >> refined over time. >> > Just like what librarians have done with the MARC data set over the > last 40 years. > > We aren't going to escape who we are and sure, our views of what MARC data means has changed over time. Why is this a problem? Understand it is a mapping problem but why does it bother you? Do you really think terms have some fixed meaning? (Reading Berlo's "The Process of Communication" where he argues that words have no meaning, only users do.) >>> Yet the library world is mostly void of understanding Topic >>> Maps, little less implementation of it. I have many online friends in >>> the library world, and they are all as frustrated as me with the lack >>> of a MARC cleanup job that might enable this "many access points" >>> dream. If there *was* such interest I know a handful of very smart >>> people who would jump on it straight away! Alas, there is a distinct >>> disjoint between what the library needs to do and the management that >>> runs it. >>> >> You mean the lack of someone else to clean up the data to our liking. >> > Why is this about "someone else"? I'm talking about cleaning up > library meta data for the benefit of librarians. The world don't > really care all that much about this, and certainly don't understand > the issues involved, probably won't cry if the library world > disappears alltogether (to them, it's the knowledge which is > important, not the box it comes in or the place they gain it). The > library needs to take some responsibility for their own meta data, > make it ready for the world, otherwise the world will simply go on > making its own and ignoring what the library has spent 200 years > building up. I would think this was terribly clear cut. > > Sure it is. The library world won't do what you want, at their expense. That part I got. You are still framing it as all or nothing. To take a somewhat unrelated example, the Gutenberg project (for all of my disagreements with it on markup issues), has in fact produced a lot of work relying solely on volunteers. But I assume the counter-argument is going to be that volunteers will lack the expertise. Sure, like all the people who graduate with degrees in biblical Greek every year lack the expertise to read NT manuscripts. Sure, there are parts that take years of training and experience, not to mention working directly with the original mss. But, not as many parts as you would think. Most mss. transcription projects are to employ graduate students. That is a world that could be opened up to the public. Would the quality vary? Sure, perhaps a bit more than by true "professionals" but look at all the interest it would generate. >> But, what if everyone cleaned up the part of greatest interest to them? I >> spend a lot of time with older CS literature so I might try my hand at some >> of those records. Other people have other areas of interest. >> > I have myself cleaned up tons of MARC regarding early music (and > specifically the context around Claudio Monteverdi), but it's a futile > exercise in the long run because new meta data suffers from the > previous errors. If you're planning on continuous integration of MARC > meta data you need a framework for filtering, matching, killing, > fixing and handling the meta data. The stupid thing is that any major > library institution, private or public, in the world!!! has got the > equivalent of this already in place, OCLC, LOC, national libraries > around the world (and I've got special knowledge of the Libraries > Australia), all huge filtering systems costing millions in people and > resources. Why aren't these efforts simply open sourced? Why aren't > these projects openly discussed, shared and extended? > > I don't know, why don't we ask them? Nicely. >> Rather than seeing the problem as an all or nothing Mount Everest of data >> conversion, all I am suggesting is that as people use a topic map based on >> such data, some of them would have enough interest to improve the topic map >> by contributing to it. Whether than would be enough or not, I honestly don't >> know. >> > Any collection of MARC you fix up will be doomed to live outside the > original context for its useful lifetime. I think you gravely > underestimate the problem of MARC, but I appreciate your positive > enthusiasm. :) I was once there, too. > > Still persisting in this your data/my data paradigm aren't you? Why would I need to fix it up outside of its original context? What I suggested was making a topic map of the data *as is* and then adding more mappings to the topic map. People who are happy with the original data that you deride so much are free to continue to use it. People who want to improve that same data can do so. People who want to use the improved data can do so. What am I being unclear about? Not to mention that if I were to use one of Robert Barta's virtual topic maps, I don't even have to copy the data at all. ;-) >> I do know that waiting for a library messiah to come along and convert >> decades worth of data to some unspecified level of quality is even less >> likely to be effective. >> > I've given up on the library world. They will not be able to save > themselves from going under, and the library will slowly turn into an > archive of objects of peripheral interest. But that's just little > positive me talking. > > On the contrary, like people in the SW echo chamber, I think you are mistaking your experiences for the world. I know any number of people who are working to make a positive difference in the library community. Granted there are probably even more people trying to making sure there is no change at all or what change there is doesn't exceed their modest comfort levels. But how is that different from any other organization, be it banking, oil production, etc.? Unless you have located the mythic company where even mid-level management makes some contribution to the enterprise. The so-called "search" engines (a more accurate name would be "popularity" or perhaps "American Idol" engines, I guess the second one is already taken) are failing now. That is only going to get worse. The library community can offer a viable alternative to that failure. It won't be easy, or uniform but I see no reason to "give up" on the library world (or any other world for that matter). Those worlds are reflections of ourselves and I am not quite ready to give up on myself just yet. ;-) Hope you are looking forward to a great weekend! Patrick > Regards, > > Alex > -- Patrick Durusau patrick-Q/[email protected] Chair, V1 - US TAG to JTC 1/SC 34 Convener, JTC 1/SC 34/WG 3 (Topic Maps) Editor, OpenDocument Format TC (OASIS), Project Editor ISO/IEC 26300 Co-Editor, ISO/IEC 13250-1, 13250-5 (Topic Maps) Another Word For It (blog): http://tm.durusau.net Homepage: http://www.durusau.net Twitter: patrickDurusau