Re: Dynamic vs. Fixed World Views was Re: MARCXML to Topic Maps? MODS to Topic Maps?
Alexander Johannesen <[email protected]>
| Newsgroups | gmane.text.xml.xtm.general |
|---|---|
| Message-ID | <[email protected]> |
Hi, marijane white <[email protected]> wrote: > If an ILS can parse a MARC record, why can't a MARC record be parsed into a > Topic Map? It's not that it can't happen, it's just that it's a hard struggle and doesn't seem to find momentum to actually happen. > If they really are as rubbish as you claim, how is it possible > to do anything at all with them? How is it that library programmers have > created things like, say, SolrMarc, and the things built upon them like > VuFind and Blacklight? All of these projects do pretty much the same thing ; export out of your ILS / OPAC your MARC records, filter and / or clean the data, recreate the MARC model to various degrees, and then isolate the new data set to make use of it. At no point does the cleaned data get back into the MARC infra-structure because there is no where for them to go. The "rubbish" statement picks more at two specific things with the MARC data more than anything ; 1. Little to no typed data, no enumeration of values 2. Little to no identity management Then there's the larger problem of the MARC infra-structure (the library infra-structure) where the tools that deal with MARC have no validation of values. This is why I'm so irate about MARCXML because they claim one of the big advantages is validation, but you can only validate that the wrapper is sane, not that the content is. And the content is built by hand (mostly) by librarians for over 40 years. The infra-structure and the tools do not and can not validate that fields that are supposed to contain dates or date ranges indeed contain dates or date ranges, cannot validate that a field which contain, say, a reference to an author indeed points to an author. Here's the problem a bit more specified ; Every field (and I include subfields and indicators here) has a name, and in the MARC21 standard has a host of rules and best-practices. However, the rules are very often not something that can be automatically checked. For example, a field that contains date range might have "19th - 20th", or "ca. 3 bc.", "1910-1940", "50's", "classical period" in them. It's well and fine to write special parsing rules to extract roughly what is meant and then represent that in some ISO format, but remember that there are thousands of fields, all different in what they represent, and often you'll find that a record represents an item with a specific compound of fields (so that if this field and that field says X and Z, then if field A is this, then it's a paperback, if not it's a hardcopy, and so on) making it very hard for typification. (I tried to find one of these official OCLC or LOC documents that lists the process for just determining what type of thing any given record is, but my Google-foo is weak this morning. Suffice to say, it's a fun process!) Basically MARC is hundreds of free-text fields that needs to be parsed and checked against some back-end facit. As stated earlier, all large libraries and organisations *already* have their own projects in place that do this kind of thing (for a variety of reasons, too, like they want to clean up the MARC, or for outside import / export reasons), from NLA (with Libraries Australia), OCLC, LOC, other national libraries, Library Thing, OpenLibrary, and on and on. But we don't have one we all can use and collaborate on. What should happen is of course a project that tries to do all this, an open-source project all of us should join in order to create this huge mapping, be it ontologically or software or processes and filter specifications, or whatever, I'm not fussy, anything will do. But it doesn't exist, nor does it look like it will be created as too many librarians don't see the problem, and outsiders will look at it and probably not think it worth their time. I shudder at the thought of all those resources wasted because of this situation. If we all value the MARC data set so much, and we all agree that it should be free (to which they all don't agree) it should be done, methinks, but then a lot of people seem to disagree with me on this. :) > Is the situation really so dire? I think so. As the library world is becoming less and less valuable to society at large, the need for their MARC data set decrease as outside bodies are finding other alternatives instead as the library world can't or won't be more open about this. Knowledge is moving out of books and into electronic form where non-librarians can do a better job at sifting through information in search of the nuggets. The bibliographic world will slowly turn into museums of special interest objects. But I don't know how long this will take. Don't get me wrong, there's always going to be libraries, but they will not be the bastion of knowledge as they have been for centuries, and it will be very interesting to watch and see what sacrifices and compromises that will happen. A lot of librarian values will be lost. Anyway, that's my positive pep-talk for the day. Regards, Alex -- Project Wrangler, SOA, Information Alchemist, UX, RESTafarian, Topic Maps --- http://shelter.nu/blog/ ---------------------------------------------- ------------------ http://www.google.com/profiles/alexander.johannesen ---