re: Invalid UTF-8 characters causing MARC::Record crash.

[email protected] (Edmund Chamberlain) Fri, 17 Jun 2011 10:53:00 +0100
Newsgroups perl.perl4lib
Message-ID <[email protected]>
Firstly, hello! Its my first time posting and possibly somewhat 
predictably with a call for help with Unicode stuff.

I've just checked the archive and seen this thread and am having a 
similar problem, a badly encoded character is causing a while loop 
through MARC::Batch->next to crash out with:

utf8 "\x87" does not map to Unicode at 
/usr/lib64/perl5/5.8.8/x86_64-linux-thread-multi/Encode.pm line 173.

I've tried pasting Al's modified decode subroutine and package to the 
script, but it is still failing. One of the offending records is 
isolated and attached.

Any suggestions welcome with regards to further modifying the sub or 
alternatives to MARC::Batch->next would be welcome. For the scope of the 
project, I'm limited to large batch files of Marc21.

With thanks

Ed



-- 
Edmund Chamberlain
Systems Development Librarian	
Electronic Services and Systems
Cambridge University Library
West Road,
Cambridge
CB3 9DR

tel: (+44) 01223 747437
fax: (+44) 01223 333160

email: [email protected]

Try LibrarySearch at http://search.lib.cam.ac.uk - a new way to discover 
Cambridge Library Collections
dodgy_char.mrc (application/octet-stream, 490 B) - not displayed