Monitoring

Buchan Milne <[email protected]> Tue, 20 Dec 2005 10:59:53 +0200
Newsgroups gmane.linux.mandrake.server
Message-ID <[email protected]>
--nextPart1712813.d0B0gM1Yf6
Content-Type: text/plain;
  charset="us-ascii"
Content-Transfer-Encoding: quoted-printable
Content-Disposition: inline

I wonder if it would be useful discusssing monitoring, and which monitoring=
=20
tools we should concentrate more on, and possibly try and get one of them=20
into main (it seems weird that we have no monitoring tool in main ...).

As some of you know, I currently work for an ISP. We have two separate=20
deployments, with about 35 production servers running RHEL in one, and abou=
t=20
16 production servers running RHEL in the other. We deploy the servers usin=
g=20
kickstart files generated from a configuration database, so that all=20
packages, configurations etc for the host are done during installation (whi=
ch=20
in the case of machines with lights-out management can be done/initiated fr=
om=20
anywhere that has network access).

(A number of packages we run on RHEL are rebuilds of Mandriva packages,=20
including a number I maintain, such as OpenLDAP, hobbit etc).

We are currently running Nagios as the monitoring tool for the larger=20
deployment, but due to a number of reasons, we have decided to run Hobbit (=
a=20
BigBrother clone) as the monitoring tool for the smaller deployment, and ha=
ve=20
also set Hobbit up for the larger deployment. This (non-production) Hobbit=
=20
installation for the larger deployment is accessible at present, at=20
http://196.25.211.20/hobbit/ .

Anyway, some of the reasons we are using Hobbit are:

=2Ddoes both status monitoring/alerting and trend monitoring
=2Dintegrated trend monitoring of all native checks (disk, cpu, memory, tcp=
=20
service response times)
=2Dbuilt-in ssl certificate checks for any ssl-enabled service (ie https)
=2Da large collection of additional checks (from http://www.deadcat.net),=20
including all BigBrother extension scripts. For example, we monitor the HP=
=20
Insight Manager snmp data via the CIM extension).
=2Dease of writing extensions, and being able to monitor trends in the resu=
lts=20
of those extensions (via the ncv rrd plugin), for example:=20
http://196.25.211.20/hobbit-cgi/bb-hostsvc.sh?HOSTSVC=3Dio.ol which uses th=
is=20
script I wrote: http://www.zarb.org/~bgmilne/bb-openldap.pl
=2Dless complexity in the interface than Nagios, while supporting (AFAICS) =
all=20
the functionality
=2Dless complexity in configuration (ie configuration for service monitorin=
g for=20
each hosts consists of one line in the bb-hosts file)


Other tools that I haven't really had a chance to look at in much detail:

1)monitorix, which Antoine uploaded recently
http://www.monitorix.org
Seems to have most of the features of Hobbit.

2)Oreon
http://www.oreon-project.org/

Seems to be a better Nagios frontend, addresses some of the problems I have=
=20
with Nagios.

3)Zabbix
http://www.zabbix.org/

Also has win32 clients.

4)Cacti
AFAIK, just trend monitoring using rrd+MySQL

So, should we package all the ones we don't have yet, or decide on features=
=20
that are necessary, and package/support only one or two?

Regards,
Buchan

=2D-=20
Buchan Milne
B.Eng,RHCE(803004789010797),LPIC-2(LPI000074592)

--nextPart1712813.d0B0gM1Yf6
Content-Type: application/pgp-signature

-----BEGIN PGP SIGNATURE-----
Version: GnuPG v1.4.2 (GNU/Linux)

iD8DBQBDp8gPrJK6UGDSBKcRAuAIAJ9otMlEPeqfnxzo3USr+JBU0m07YACbBiGw
9fWA2iIXc8qJpd4ifUezfhM=
=Dh2v
-----END PGP SIGNATURE-----

--nextPart1712813.d0B0gM1Yf6--