LibXML: Ghost attributes
Daniel Perrett <[email protected]>
| Newsgroups | gmane.comp.lang.perl.xml |
|---|---|
| Message-ID | <OFBE1EC97D.C22EFED9-ON80257575.005E1139-80257576.006292E8@cambridge.org> |
Hello
I'm using LibXML to do some processing on a directory containing some
highly structured dictionary data (some files have over 15k elements, and
most elements have one or more attributes), outputting to another folder.
I'm finding that nodes from early files start creeping into files I
process later (and possibly also the same file). This happens less often
when I create a new parser for each file, but it's not resolved.
Specifically, I am adding new attributes to newly created elements in one
file, and those attributes with those values are also appearing when I run
the same process on another file.
The full code is very long, but below are some relevant bits, edited for a
little readability. It's also not particularly transparent what I'm doing
unless you already know the structure of the dictionary well. Basically,
the chunk I've isolated inserts a new element, <infs> into every <sense>
within an <entry>, which has attributes @id and possibly @var, the first
of which reads something like daily*adverb, the second is a comma
separated list of all the variants. Thing is, I occasionally get things
like <infs id="date back*verb" var="to judge by/from">...</infs> (in this
case, I've processed jk.xml (containing variant "to judge by/from") before
us_d.xml).
Is there any other housekeeping I should be doing once I get to the end of
a file or while I'm working on it? Are the references to the senses
colliding or something? Or is it something else more simple?
Apologies for any bad practices, I've only been using perl for 6 months or
so. Indentation is a bit out, due to copy-pasting.
Daniel
#####
# The paths that are of interest are:
# epp/entry/posblock/gwblock/phraserec/phrase
# the phrase that 'heads up' the phraserec
# epp/entry/posblock/gwblock/phraserec/var/v
# a variant form of that phrase
# epp/entry/posblock/gwblock/phraserec/sense
# one 'sense' of a phrase (in this dictionary, each phrase has only
one).
# epp/entry/posblock/gwblock/phraserec/sense/infs
# not in the original data, but something I create in the course of
the perl script.
# ... where epp is the root element.
my %sVariants; # each value, keyed to the sense element, is an array of
strings; each string is a variant of the sense element.
sub proc_file($) # called on each xml file in a directory.
{
my $file=shift;
my $oNewParser = XML::LibXML->new;
$oNewParser->keep_blanks(0);
my $doc = $oNewParser->parse_file($file);
my $eRoot = $doc->documentElement();
foreach my $eSense ($eRoot->findnodes('.//sense'))
{
proc_sense($eSense);
}
}
sub proc_sense($)
{
my $eSense = shift;
my $eEntry = get_anc($eSense, 'entry');
my $isPhrase;
my $eSW;
# I declare all my variables, I'm using strict, I'm sure I'll have
missed something in the lines I've copied out, but strict is happy with my
code.
# eSense is the <sense> element, eEntry the entry element.
# get_anc is a function which returns the first element going up
the tree that is an ancestor of the first argument with the node name in
the second argument
# Some senses are inside a phrase record element. We add
infs wherever this is true.
if (defined ( get_anc($eSense, 'phraserec')))
{
my $ePhraserec=get_anc($eSense, 'phraserec');
$eSW = ${[$ePhraserec->findnodes('./phrase')]}[0];
# there's only ever one of them
if (defined ($eSW)) # just in case
{
$sSW = strip_phrase($eSW); # strip_phrase
reduces a phrase to text content and removes certain child elements
$isPhrase=1;
}
foreach ($ePhraserec->findnodes('./*/v')) # get
all the variants in the phraserec
{
push @{$sVariants{$eSense}},
strip_phrase($_); # add the text in the variant to a list of them which is
unique to this sense
}
}
# I also get variants from other paths in the xml, which I
haven't repeated. But these other places are all within the same xml file,
while some of the errors I get come from other xml files.
my $eSWinfs = $eSense->addNewChild('', 'infs');
my $sVarList = "";
if (defined (${$sVariants{$eSense}}[0]))
{
$sVarList = join (',',
@{$sVariants{$eSense}});
}
if ($sSW =~ m/\w/)
{
$eSWinfs->setAttribute('id',"$sSW\*$sPosList"); # poslist is determined
elsewhere within proc_sense, I've not copied out the code.
$eSWinfs->setAttribute('var',"$sVarList")
if $sVarList; # offending attribute
if ($isPhrase)
{
$eSWinfs->setAttribute('id',"$sSW\*phrase");
}
}
}
_______________________________________________
Perl-XML mailing list
[email protected]
To unsubscribe: http://listserv.ActiveState.com/mailman/mysubs