Re: Asunto: Re: Slow fork bomb message in latest version of POE

Rocco Caputo <[email protected]> Mon, 24 Mar 2014 13:30:39 -0400
Newsgroups gmane.comp.lang.perl.poe
Message-ID <[email protected]>
Hi again.

What I mean is that I don't think you know what sig_child() does =
exactly, or how to use it.  I base this impression on two things: First, =
you're calling sig_child() from a place where it will never work and at =
a time that is obviously too late to do anything.  Second, it needs at =
least two parameters to work, but you're passing it nothing.

I recommend not using SIGDIE for common exception handling.  Its scope =
is too broad, and your code will get ugly.  It's probably cleaner to use =
eval{} or Try::Tiny to convert your unexpected exceptions into expected =
ones.  If you catch them explicitly, then POE won't need to raise them, =
and there should be less strange behavior.

The problem seems to be migrating.  I recommend caution against further =
clouding the original issue until it's resolved.

If you resolve your exceptions issue, and if you resolve your =
sig_child() usage issue, then your program should not be interrupted at =
inopportune times, and it should reap the nginx process before it exits. =
 This should resolve all outstanding issues, as I currently understand =
them.

--=20
Rocco Caputo <[email protected]>

On Mar 24, 2014, at 12:15, albertocurro <[email protected]> wrote:

>=20
> Hi,
>=20
> Sorry, but I don't catch what you exactly mean with "not using =
sig_child() as intended". Do you mean calling it from the main session =
so each child process will be closed properly?=20
>=20
> The issue I have is how to handle unexpected exceptions. Seems they =
are thrown and raised without control, killing POE's kernel before in =
the way. I could be thinking in the timing in the wrong way, though...
>=20
> Alberto
>=20
> ---- Activado lun, 24 mar 2014 16:59:49 +0100 Rocco =
Caputo<[email protected]> escribi=F3 ----=20
>=20
>> You are not using sig_child() as intended.  When used as intended, =
sig_child() will prevent shutdown until the child process has exited and =
has been reaped.  The timing issues you're worried about should not =
exist.=20
>>=20
>> -- =20
>> Rocco Caputo <[email protected]>=20
>>=20
>> On Mar 24, 2014, at 11:44, albertocurro <[email protected]> =
wrote:=20
>>=20
>>> Hi Rocco,=20
>>>=20
>>> many thanks for your quick answer! Unfortunately, the provided =
solution only works partially. I still have some cases where the "fork =
bomb" message is here with us :(=20
>>>=20
>>> One of the cases is this one: under some configuration, an instance =
of nginx is started, so our product writes the configuration file and =
starts the Nginx instance pointing to that configuration file. BUT, if =
the configuration file could not be written (directory does not exist, =
etc), then the error raises, and I've not found any way to handle it:=20
>>>=20
>>> DEBUG - Created nginx temporary directory /opt/tmp/pull/instance1=20
>>> DEBUG - Created nginx configuration directory =
/opt/etc/pull/instance1=20
>>> DEBUG - Created nginx log directory /opt/log/pull/instance1=20
>>> DEBUG - creating nginx configfile for instance 1 in =
/opt/etc/pull/instance1=20
>>> =3D=3D=3D 13991 =3D=3D=3D !!! Kernel has 1 child process(es).=20
>>> =3D=3D=3D 13991 =3D=3D=3D !!! At least one child process is still =
running when POE::Kernel->run() is ready to return.=20
>>> =3D=3D=3D 13991 =3D=3D=3D !!! Be sure to use sig_child() to reap =
child processes.=20
>>> =3D=3D=3D 13991 =3D=3D=3D !!! In extreme cases, failure to reap =
child processes has=20
>>> =3D=3D=3D 13991 =3D=3D=3D !!! resulted in a slow 'fork bomb' that =
has halted systems.=20
>>> Could not open file: No such file or directory=20
>>>=20
>>> I've added a DIE handler in the main session to try to handle this:=20=

>>>=20
>>> $sig_session =3D POE::Session->create(=20
>>>   inline_states =3D> {=20
>>>       _start =3D> sub {=20
>>>           $_[HEAP]{RELOADED} =3D 0;=20
>>>           $_[KERNEL]->sig(TERM =3D> '_sigterm');=20
>>>           $_[KERNEL]->sig(INT =3D> '_sigterm');=20
>>>           $_[KERNEL]->sig(DIE =3D> '_sigterm');=20
>>>           $_[KERNEL]->sig(nginx_reload =3D> '_sig_nginx_reload');=20
>>>           $_[KERNEL]->alias_set('sighandler');=20
>>>       },=20
>>>       _sigdie =3D> sub {=20
>>>           print "Handling exception, calling stop";=20
>>>           POE::Kernel->call($sig_session, '_stop');=20
>>>       },=20
>>>       _stop =3D> sub {=20
>>>           # Reap any existing pid (# 1825119)=20
>>>           print "Handling stop";=20
>>>           POE::Kernel->sig_child();=20
>>>           use POSIX ":sys_wait_h";=20
>>>           1 while waitpid(WNOHANG, -1) > 0;=20
>>>=20
>>>           # Clear signal handlers...=20
>>>           $_[KERNEL]->sig('TERM');=20
>>>=20
>>> But, as said above, it's not working. Checking POE's code, I can see =
the message lines are generated in Resources/Signals.pm, under =
_data_sig_finalize() method (where POE is already doing the same you =
recommended me, waiting for the pid).=20
>>>=20
>>> But _data_sig_finalize() method is called in Kernel.pm just after =
unregistered all the signals (Kernel.pm =3D> _finalize_kernel):=20
>>>=20
>>> my $self =3D shift;=20
>>>=20
>>> # Disable signal watching since there's now no place for them to go.=20=

>>> foreach ($self->_data_sig_get_safe_signals()) {=20
>>>   $self->loop_ignore_signal($_);=20
>>> }=20
>>>=20
>>> # Remove the kernel session's signal watcher.=20
>>> $self->_data_sig_remove($self->ID, "IDLE");=20
>>>=20
>>> # The main loop is done, no matter which event library ran it.=20
>>> # sig before loop so that it clears the signal_pipe file handler=20
>>> $self->_data_sig_finalize();=20
>>> $self->loop_finalize();=20
>>>=20
>>> Once here, none of my signal handlers in the main session instance =
would work, as the signals have been unregistered. On an exception (die) =
while POE::Kernel->run(), how could I handle it then??=20
>>>=20
>>> Thanks a lot=20
>>> Alberto=20
>>>=20
>>>=20
>>>=20
>>>=20
>>> ---- Activado lun, 24 mar 2014 13:45:45 +0100 Rocco Caputo  escribi=F3=
 ---- =20
>>>=20
>>>> Hi, Alberto. =20
>>>>=20
>>>> At program end time, POE runs a quick waitpid() check for child =
processes that may have leaked. This check was added after a bug report =
where POE locked up a server after several days of running. It turned =
out to be the reporter's application, but it was hard to debug. =20
>>>>=20
>>>> Your program seems to have created two processes that it didn't =
reap: PIDs 5373 and 5374. The ideal solution is to reap those processes =
before exiting. Your program can do this using POE::Kernel's sig_child() =
method. =20
>>>>=20
>>>> In some cases, a third-party library will create processes and not =
properly clean them up. It can be impossible to solve this case without =
modifying other people's code. =20
>>>>=20
>>>> If you just want to ignore the problem, this might do the trick. =
Put these lines in your last _stop handler. They should reap the =
processes you've leaked before POE's check: =20
>>>>=20
>>>> use POSIX ":sys_wait_h"; =20
>>>> 1 while waitpid(WNOHANG, -1) > 0; =20
>>>>=20
>>>> It's a bit of a pain, but I think it's better to explicitly ignore =
the problem than for it to go unnoticed by default. =20
>>>>=20
>>>> Please let me know whether that resolves your problem. It may not. =
For example, the processes may still be open until an object is =
destroyed at global destruction time. =20
>>>>=20
>>>> -- =20
>>>> Rocco Caputo  =20
>>>>=20
>>>> On Mar 24, 2014, at 05:46, albertocurro  wrote: =20
>>>>=20
>>>>> Guys, =20
>>>>>=20
>>>>> We have a product developed using POE as a base framework, with =
some other tool libraries as log4perl; basically is a forward proxy, =
composed of several modules, each one of them comprising a POE::Session; =
all of them share an internal queue of tasks to be performed. Each =
module performs several tasks on initialization, and if anything goes =
wrong, croak() is called to stop the service -> this is considered ok, =
since croak() is only called during initialization, when validation is =
being performed. =20
>>>>>=20
>>>>> The product is stable and works really fine, but recently I =
updated POE to the latest version, and since then we can see this =
message in the logs: =20
>>>>>=20
>>>>> registering pdu failed: 263! =20
>>>>> =3D=3D=3D 5267 =3D=3D=3D 5 -> on_handle (from =
Handler/StoreRemote.pm at 87) =20
>>>>> =3D=3D=3D 5267 =3D=3D=3D 5 -> on_retry (from =
Handler/StoreRemote.pm at 141) =20
>>>>> =3D=3D=3D 5267 =3D=3D=3D 9 -> on_handle (from =
Handler/StoreRemote.pm at 87) =20
>>>>> =3D=3D=3D 5267 =3D=3D=3D 9 -> on_retry (from =
Handler/StoreRemote.pm at 141) =20
>>>>> =3D=3D=3D 5267 =3D=3D=3D !!! Kernel has child processes. =20
>>>>> =3D=3D=3D 5267 =3D=3D=3D !!! Stopped child process (PID 5373) =
reaped when POE::Kernel->run() is ready to return. =20
>>>>> =3D=3D=3D 5267 =3D=3D=3D !!! Stopped child process (PID 5374) =
reaped when POE::Kernel->run() is ready to return. =20
>>>>> =3D=3D=3D 5267 =3D=3D=3D !!! At least one child process is still =
running when POE::Kernel->run() is ready to return. =20
>>>>> =3D=3D=3D 5267 =3D=3D=3D !!! Be sure to use sig_child() to reap =
child processes. =20
>>>>> =3D=3D=3D 5267 =3D=3D=3D !!! In extreme cases, failure to reap =
child processes has =20
>>>>> =3D=3D=3D 5267 =3D=3D=3D !!! resulted in a slow 'fork bomb' that =
has halted systems. =20
>>>>> mkdir /mnt/nfs99: Permission denied at Handler/Store.pm line 147 =20=

>>>>>=20
>>>>> first lines and last line above are the errors itself, but this =
part is new since the upgrading: =20
>>>>>=20
>>>>> =3D=3D=3D 5267 =3D=3D=3D !!! Kernel has child processes. =20
>>>>> =3D=3D=3D 5267 =3D=3D=3D !!! Stopped child process (PID 5373) =
reaped when POE::Kernel->run() is ready to return. =20
>>>>> =3D=3D=3D 5267 =3D=3D=3D !!! Stopped child process (PID 5374) =
reaped when POE::Kernel->run() is ready to return. =20
>>>>> =3D=3D=3D 5267 =3D=3D=3D !!! At least one child process is still =
running when POE::Kernel->run() is ready to return. =20
>>>>> =3D=3D=3D 5267 =3D=3D=3D !!! Be sure to use sig_child() to reap =
child processes. =20
>>>>> =3D=3D=3D 5267 =3D=3D=3D !!! In extreme cases, failure to reap =
child processes has =20
>>>>> =3D=3D=3D 5267 =3D=3D=3D !!! resulted in a slow 'fork bomb' that =
has halted systems. =20
>>>>>=20
>>>>> I can see it everytime the service is stopped because of an =
unhandled condition, even when POE's event loop has been already running =
for ours. It was not visible before, and I can't get rid of it in any =
way. I've tried different ways to avoid it with no luck. =20
>>>>>=20
>>>>> Any advice or alternative approach on this? =20
>>>>>=20
>>>>> Many thanks =20
>>>>> Alberto=20
>>=20
>>=20
>=20