RE: Is this best for multiline parsing?

[email protected] ("Orton, Yves")
Newsgroups perl.recdescent
Message-ID <71B318898201D311845C0008C75DAD1C063A271B@defra1ex2>
While I think this is interesting, (I believe Randal Schwartz did a column
on doing something similer,) it seems to me that MIME::Parser already has
the functionality that you want. Perhaps you should look there?

Yves

> -----Original Message-----
> From: Chris Wilkes [mailto:[email protected]]
> Sent: 2002/02/19 11:11
> To: [email protected]
> Subject: Is this best for multiline parsing?
> 
> 
> Hello,
> 
>   I'm working on Mail::Audit so that I can pull out file attachment
> names from it.  To do so requires that you look at the headers for a
> "boundary" marker to figure out what seperates different attachments.
>   This is my first attempt at using RecDescent and would like 
> to know if
> this looks like the best way to do it or if someone would offer some
> style tips.
>   My main inspiration in this script was Damian's post of how 
> to parse a
> command seperated over multiple lines.  Except that one was a little
> easier than this as it had a line continuation marker at the end.
> Looking through some mail headers they don't do this, rather just
> indenting the next line a couple of spaces.
>   My next question is how to continue on with the rest of the file.  I
> don't really care what the file attachment is beyond its 
> listed filename
> and what lines it takes up in the ->body() reference to an 
> array returned
> by Mail::Audit.  I'm thinking about doing a while() loop till I get to
> the boundary marker and then sending the next non-blank lines through
> another RecDescent parse.  I don't want to have to tell RecDescent to
> look through thousands of lines of a mail file when it is not needed.
>   I realize that most of this could be done without RecDescent but
> thought I would give it a try.  Maybe I'm a masochist :)
> 
> Chris
> 
> ----------cut----------cut----------cut----------cut----------
> #!/usr/bin/perl
>  
> use Data::Dumper;
> use Parse::RecDescent;
> use strict;
>  
> my ($result, $parser, $grammar, $header);
>  
> $header = getdata();
>  
> $grammar = join '', <DATA>;
> $parser = Parse::RecDescent->new($grammar)
>   or die "Bad grammar '$grammar'\n";
>  
> $result = $parser->headerinfo($header);
>  
> print Dumper($result);
>  
> sub getdata {
>   my $header = <<'EOF';
> Delivered-To: [email protected]
> Content-Type: Multipart/Mixed;
>   charset="iso-8859-1";
>   boundary="------------Boundary-00=_2YQQSZM4Z8EP8VA9F3K8"
> From: Joe Blow <[email protected]>
> EOF
>   return $header;
> }
>  
> __DATA__
>   headerinfo:
>     parameter(s)
>     {  my %return;
>        foreach (@{$item{parameter}}) {
>          push @{$return{$_->[0]}}, $_->[1];
>        }
>        { \%return } }
>   parameter: key ':' value
>       { [ $item{key}, $item{value} ] }
>   key:
>     /^([^:]+)/ { $1 }
>   value:
>     multiline_value | singleline_value
>     singleline_value:
>       /(.*)\n/ { $1 }
>     multiline_value:
>       /^(.*?)\s*\n(\s+.*\n)+/
>         { $item[-1] =~ s/^\s+/ /gm;
>           $item[-1] =~ s/\s*\n//gm;
>           { $item[-1] } }
>
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.