Re: looping 2 times through 5000 differs from 5000 times through 2

[email protected] (oldgeezer)
Newsgroups perl.beginners
Organization http://groups.google.com
Message-ID <[email protected]>
On Jun 8, 5:03 pm, [email protected] (Rob Dixon) wrote:
> oldgeezer wrote:
> > Hi all,
>
> > Last week I discovered this perl.beginners group.
> > Good stuff here, albeit many times hard to grasp
> > the answers. But I'm learning.
>
> > What I would like to understand is why looping
> > 2 times through 5000 lines takes less time than
> > looping 5000 times through 2 lines.
>
> > To show what I mean, I wrote a snippet that
> > does nothing with the data and yet the first
> > part is 5 times faster than the second part
> > (in my pc anyway).
>
> > Thanks.
> > Rob.
>
> > #!/usr/bin/perl
>
> > use strict;
> > use warnings;
> > use POSIX qw(clock);
>
> > #just create a small array for this test
> > my @few=();
> > for my $i(1...2){push @few, $i};
>
> > #just create a big array for this test
> > my @many=();
> > for my $i(1...500000){push @many, $i};
>
> > # part 1, loop few times over much data
> > my $time_1=clock();
> > foreach my $val (@few){
> >     foreach my $other(@many) {} #do zilch
> > }
> > my $time_2=clock();
> > my $delta1=$time_2-$time_1;
> > print "Delta1 $delta1\n";
>
> > # part 2, loop many times over few data
> > my $time_3=clock();
> > foreach my $val (@many){
> >     foreach my $other(@few) {} #do zilch
> > }
> > my $time_4=clock();
> > my $delta2=$time_4-$time_3;
> > print "Delta2 $delta2\n";
>
> > exit 0;
>
> In addition to what others have said, it is worth doing the same experiment for
> an empty loop:
>
>   foreach my $val (@many){
>   }
>
> to see how much overhead there is for each iteration of the loop.
>
> You should also investigate the Benchmark module, which makes comparisons such
> as these more convenient.
>
> Rob

Hi all,

Sorry for the late response, I've been busy doing something else.

As a newbie in this perl.beginners, I did not know that it is a
good thing to tell what purpose the script has.
My script is purely hobby. All my scripts are.

I understand it now, I simply overlooked:
>> How many times do you create a new iteration
>> variable <snip>

And then someone replied:
>>In either case ... do not worry.
I did not worry. My script works okay,
I only wondered why reshuffling a loop
made a big difference in time.

I wasn't aware that a  Benchmark module exists.
Seems worth playing with it. Which I will do.

The code I wrote in my question was only to give
you an idea what I was talking about, not the real
code because then I would have given the data
that I work on as well.
It is not secret data, but they are long lines and I
was afraid they would wrap.

But you also made me aware of a mistake
I always made until now. The three periods in
     for (x...y)
I must have inherited that error from
another language. Probably from an
interpreter I wrote myself some 30
years ago.
Thanks for all the help.

Below I give you the original lines in my script.
I usually write a lot of comment lines, because
I know myself. After some time I forget why I did
something, and then reinvent the wheel.
The script is a CGI. Data in the data base
is separated with a  ';'  (is that called a semi-colon?
I am not native english).

foreach my $entry (uniq @IK){
  foreach my $line(@DATA) {
    # Note /;$entry;/ takes longer than
    # /$entry;/ and that takes longer than /$entry/
    # But /$entry/ also splits remarks_entries.
    # So I need the trailing ';' because they
    # are not in the remarks_entries
    # Appending the i, like: /$entry/i slows
    # down too. I don't need it.
    # Prepending m, like m/$entry/
    # does not change anything (timewise).
    # For the average person /$entry;/ is fastest
    if($line=~ /$entry;/) {
     splitline($line); #@IK changes when I splitline
     #next; don't do this. It makes it slower
    };
  };
};

Swapping the first two foreach-lines make a
big time difference, @DATA is the data base
and is about 5000 lines, @IK varies from 1
to around 20 strings, dependent on the request
POSTed to the CGI-script.

'uniq' and 'splitline' are subs. There is
a trick. @IK changes in the loop, but
the 'uniq @IK' is only defined once and
is not changed by the loop. The changed @IK
is used later again for some other reason.

Oh, and those trailing ';' after all the '}'
may well be superfluous. But I like to see
when a } is really the end of a block
and will not be followed by an else or
something like that. I know: emacs does
a nice indentation, but still I (personally)
like to see that ';'

Thanks again,

Rob.
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.