Re: Splitting a large file of MARC records into smaller files

[email protected] (Ashley Sanders)
Newsgroups perl.perl4lib
Message-ID <[email protected]>
Jennifer,

> I am working with files of MARC records that are over a million records each. I'd like to split them down into smaller chunks, preferably using a command line. MARCedit works, but is slow and made for the desktop. I've looked around and haven't found anything truly useful- Endeavor's MARCsplit comes close but doesn't separate files into even numbers, only by matching criteria, so there could be lots of record duplication between files.
>
> Any idea where to begin? I am a (super) novice Perl person.

Well... if you have a *nix style command line and the usual
utilities and your file of MARC records is in exchange format
with the records just delimited by the end-of-record character
0x1d, then you could do something like this:

tr '\035' '\n' < my-marc-file.mrc > recs.txt
split -1000 recs.txt

The tr command will turn the MARC end-of-record characters
into newlines. Then use the split command to carve up
the output of tr into files of 1000 records.

You then may have to use tr to convert the newlines back
to MARC end-of-record characters.

Ashley.

-- 
Ashley Sanders               [email protected]
Copac http://copac.ac.uk A Mimas service funded by JISC
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.