StrongHold trouble

Ricardo Manuel Oliveira <[email protected]> Thu, 03 Mar 2005 17:30:00 +0000
Newsgroups gmane.linux.redhat.stronghold
Message-ID <[email protected]>
Hi,

  I have a server farm with a couple StrongHold 3 and StrongHold4 
installations.

  Today, totally out of the blue (after about 2 to 3 years of 
uninterrupt service) the stronghold parent process died (it's a 
StrongHold4), while about 30 child processes kept running (usually we 
have about 500 to 1200 child processes running at this time). I could 
establish a TCP connection to the ports the server was listening to, but 
I'd get no response to my requests (GET, HEAD, etc).

  The only messages available in the error logs at the time are:

[Thu Mar  3 16:36:21 2005] [info] mod_unique_id: using ip addr 172.30.66.23
[Thu Mar  3 16:36:22 2005] [notice] Stronghold configured -- resuming 
normal operations
[Thu Mar  3 16:36:22 2005] [info] Server built: Oct 15 2004 12:20:49
[Thu Mar  3 16:36:22 2005] [notice] Accept mutex: sysvsem (Default: sysvsem)
[Thu Mar  3 16:36:22 2005] [alert] Child 3882 returned a Fatal error...
Apache is exiting!

  I couldn't find nothing like this problem with a plausible 
explanation, although Google suggested some mailing lists where the 
problem is referred to as "lunar ray hitting the power cable during 
solar eclipse".

  I had to kill the child processes (since a stop-server would complain 
about "no such process with pid xxxx) and issue a start-server to get it 
up and running.

  Does anyone have a clue about where the problem could be? I can follow 
clues and get to the right conclusions when pointed in the right direction.

Thanks in advance,
  Ricardo Oliveira