Asunto: Re: Slow fork bomb message in latest version of POE

albertocurro <[email protected]> Mon, 24 Mar 2014 17:15:22 +0100
Newsgroups gmane.comp.lang.perl.poe
Message-ID <[email protected]>
Hi,

 Sorry, but I don't catch what you exactly mean with "not using sig_child()=
 as intended". Do you mean calling it from the main session so each child p=
rocess will be closed properly?=20

 The issue I have is how to handle unexpected exceptions. Seems they are th=
rown and raised without control, killing POE's kernel before in the way. I =
could be thinking in the timing in the wrong way, though...

 Alberto

---- Activado lun, 24 mar 2014 16:59:49 +0100 Rocco Caputo<[email protected]=
m> escribi=C3=B3 ----=20

 > You are not using sig_child() as intended.  When used as intended, sig_c=
hild() will prevent shutdown until the child process has exited and has bee=
n reaped.  The timing issues you're worried about should not exist.=20
 > =20
 > -- =20
 > Rocco Caputo <[email protected]>=20
 > =20
 > On Mar 24, 2014, at 11:44, albertocurro <[email protected]> wrote:=
=20
 > =20
 > > Hi Rocco,=20
 > > =20
 > > many thanks for your quick answer! Unfortunately, the provided solutio=
n only works partially. I still have some cases where the "fork bomb" messa=
ge is here with us :(=20
 > > =20
 > >  One of the cases is this one: under some configuration, an instance o=
f nginx is started, so our product writes the configuration file and starts=
 the Nginx instance pointing to that configuration file. BUT, if the config=
uration file could not be written (directory does not exist, etc), then the=
 error raises, and I've not found any way to handle it:=20
 > > =20
 > > DEBUG - Created nginx temporary directory /opt/tmp/pull/instance1=20
 > > DEBUG - Created nginx configuration directory /opt/etc/pull/instance1=
=20
 > > DEBUG - Created nginx log directory /opt/log/pull/instance1=20
 > > DEBUG - creating nginx configfile for instance 1 in /opt/etc/pull/inst=
ance1=20
 > > =3D=3D=3D 13991 =3D=3D=3D !!! Kernel has 1 child process(es).=20
 > > =3D=3D=3D 13991 =3D=3D=3D !!! At least one child process is still runn=
ing when POE::Kernel->run() is ready to return.=20
 > > =3D=3D=3D 13991 =3D=3D=3D !!! Be sure to use sig_child() to reap child=
 processes.=20
 > > =3D=3D=3D 13991 =3D=3D=3D !!! In extreme cases, failure to reap child =
processes has=20
 > > =3D=3D=3D 13991 =3D=3D=3D !!! resulted in a slow 'fork bomb' that has =
halted systems.=20
 > > Could not open file: No such file or directory=20
 > > =20
 > > I've added a DIE handler in the main session to try to handle this:=20
 > > =20
 > > $sig_session =3D POE::Session->create(=20
 > >    inline_states =3D> {=20
 > >        _start =3D> sub {=20
 > >            $_[HEAP]{RELOADED} =3D 0;=20
 > >            $_[KERNEL]->sig(TERM =3D> '_sigterm');=20
 > >            $_[KERNEL]->sig(INT =3D> '_sigterm');=20
 > >            $_[KERNEL]->sig(DIE =3D> '_sigterm');=20
 > >            $_[KERNEL]->sig(nginx_reload =3D> '_sig_nginx_reload');=20
 > >            $_[KERNEL]->alias_set('sighandler');=20
 > >        },=20
 > >        _sigdie =3D> sub {=20
 > >            print "Handling exception, calling stop";=20
 > >            POE::Kernel->call($sig_session, '_stop');=20
 > >        },=20
 > >        _stop =3D> sub {=20
 > >            # Reap any existing pid (# 1825119)=20
 > >            print "Handling stop";=20
 > >            POE::Kernel->sig_child();=20
 > >            use POSIX ":sys_wait_h";=20
 > >            1 while waitpid(WNOHANG, -1) > 0;=20
 > > =20
 > >            # Clear signal handlers...=20
 > >            $_[KERNEL]->sig('TERM');=20
 > > =20
 > > But, as said above, it's not working. Checking POE's code, I can see t=
he message lines are generated in Resources/Signals.pm, under _data_sig_fin=
alize() method (where POE is already doing the same you recommended me, wai=
ting for the pid).=20
 > > =20
 > > But _data_sig_finalize() method is called in Kernel.pm just after unre=
gistered all the signals (Kernel.pm =3D> _finalize_kernel):=20
 > > =20
 > > my $self =3D shift;=20
 > > =20
 > >  # Disable signal watching since there's now no place for them to go.=
=20
 > >  foreach ($self->_data_sig_get_safe_signals()) {=20
 > >    $self->loop_ignore_signal($_);=20
 > >  }=20
 > > =20
 > >  # Remove the kernel session's signal watcher.=20
 > >  $self->_data_sig_remove($self->ID, "IDLE");=20
 > > =20
 > >  # The main loop is done, no matter which event library ran it.=20
 > >  # sig before loop so that it clears the signal_pipe file handler=20
 > >  $self->_data_sig_finalize();=20
 > >  $self->loop_finalize();=20
 > > =20
 > > Once here, none of my signal handlers in the main session instance wou=
ld work, as the signals have been unregistered. On an exception (die) while=
 POE::Kernel->run(), how could I handle it then??=20
 > > =20
 > > Thanks a lot=20
 > > Alberto=20
 > > =20
 > > =20
 > > =20
 > > =20
 > > ---- Activado lun, 24 mar 2014 13:45:45 +0100 Rocco Caputo  escribi=C3=
=B3 ---- =20
 > > =20
 > >> Hi, Alberto. =20
 > >> =20
 > >> At program end time, POE runs a quick waitpid() check for child proce=
sses that may have leaked. This check was added after a bug report where PO=
E locked up a server after several days of running. It turned out to be the=
 reporter's application, but it was hard to debug. =20
 > >> =20
 > >> Your program seems to have created two processes that it didn't reap:=
 PIDs 5373 and 5374. The ideal solution is to reap those processes before e=
xiting. Your program can do this using POE::Kernel's sig_child() method. =
=20
 > >> =20
 > >> In some cases, a third-party library will create processes and not pr=
operly clean them up. It can be impossible to solve this case without modif=
ying other people's code. =20
 > >> =20
 > >> If you just want to ignore the problem, this might do the trick. Put =
these lines in your last _stop handler. They should reap the processes you'=
ve leaked before POE's check: =20
 > >> =20
 > >> use POSIX ":sys_wait_h"; =20
 > >> 1 while waitpid(WNOHANG, -1) > 0; =20
 > >> =20
 > >> It's a bit of a pain, but I think it's better to explicitly ignore th=
e problem than for it to go unnoticed by default. =20
 > >> =20
 > >> Please let me know whether that resolves your problem. It may not. Fo=
r example, the processes may still be open until an object is destroyed at =
global destruction time. =20
 > >> =20
 > >> -- =20
 > >> Rocco Caputo  =20
 > >> =20
 > >> On Mar 24, 2014, at 05:46, albertocurro  wrote: =20
 > >> =20
 > >>> Guys, =20
 > >>> =20
 > >>> We have a product developed using POE as a base framework, with some=
 other tool libraries as log4perl; basically is a forward proxy, composed o=
f several modules, each one of them comprising a POE::Session; all of them =
share an internal queue of tasks to be performed. Each module performs seve=
ral tasks on initialization, and if anything goes wrong, croak() is called =
to stop the service -> this is considered ok, since croak() is only called =
during initialization, when validation is being performed. =20
 > >>> =20
 > >>> The product is stable and works really fine, but recently I updated =
POE to the latest version, and since then we can see this message in the lo=
gs: =20
 > >>> =20
 > >>> registering pdu failed: 263! =20
 > >>> =3D=3D=3D 5267 =3D=3D=3D 5 -> on_handle (from Handler/StoreRemote.pm=
 at 87) =20
 > >>> =3D=3D=3D 5267 =3D=3D=3D 5 -> on_retry (from Handler/StoreRemote.pm =
at 141) =20
 > >>> =3D=3D=3D 5267 =3D=3D=3D 9 -> on_handle (from Handler/StoreRemote.pm=
 at 87) =20
 > >>> =3D=3D=3D 5267 =3D=3D=3D 9 -> on_retry (from Handler/StoreRemote.pm =
at 141) =20
 > >>> =3D=3D=3D 5267 =3D=3D=3D !!! Kernel has child processes. =20
 > >>> =3D=3D=3D 5267 =3D=3D=3D !!! Stopped child process (PID 5373) reaped=
 when POE::Kernel->run() is ready to return. =20
 > >>> =3D=3D=3D 5267 =3D=3D=3D !!! Stopped child process (PID 5374) reaped=
 when POE::Kernel->run() is ready to return. =20
 > >>> =3D=3D=3D 5267 =3D=3D=3D !!! At least one child process is still run=
ning when POE::Kernel->run() is ready to return. =20
 > >>> =3D=3D=3D 5267 =3D=3D=3D !!! Be sure to use sig_child() to reap chil=
d processes. =20
 > >>> =3D=3D=3D 5267 =3D=3D=3D !!! In extreme cases, failure to reap child=
 processes has =20
 > >>> =3D=3D=3D 5267 =3D=3D=3D !!! resulted in a slow 'fork bomb' that has=
 halted systems. =20
 > >>> mkdir /mnt/nfs99: Permission denied at Handler/Store.pm line 147 =20
 > >>> =20
 > >>> first lines and last line above are the errors itself, but this part=
 is new since the upgrading: =20
 > >>> =20
 > >>> =3D=3D=3D 5267 =3D=3D=3D !!! Kernel has child processes. =20
 > >>> =3D=3D=3D 5267 =3D=3D=3D !!! Stopped child process (PID 5373) reaped=
 when POE::Kernel->run() is ready to return. =20
 > >>> =3D=3D=3D 5267 =3D=3D=3D !!! Stopped child process (PID 5374) reaped=
 when POE::Kernel->run() is ready to return. =20
 > >>> =3D=3D=3D 5267 =3D=3D=3D !!! At least one child process is still run=
ning when POE::Kernel->run() is ready to return. =20
 > >>> =3D=3D=3D 5267 =3D=3D=3D !!! Be sure to use sig_child() to reap chil=
d processes. =20
 > >>> =3D=3D=3D 5267 =3D=3D=3D !!! In extreme cases, failure to reap child=
 processes has =20
 > >>> =3D=3D=3D 5267 =3D=3D=3D !!! resulted in a slow 'fork bomb' that has=
 halted systems. =20
 > >>> =20
 > >>> I can see it everytime the service is stopped because of an unhandle=
d condition, even when POE's event loop has been already running for ours. =
It was not visible before, and I can't get rid of it in any way. I've tried=
 different ways to avoid it with no luck. =20
 > >>> =20
 > >>> Any advice or alternative approach on this? =20
 > >>> =20
 > >>> Many thanks =20
 > >>> Alberto=20
 > =20
 >=20