How to clean up outages that are not in the outage table
James Zuelow <[email protected]>
| Newsgroups | gmane.network.opennms.general |
|---|---|
| Message-ID | <[email protected]> |
I must admit that I haven't been keeping up on OpenNMS Horizon internals. I've had the same installation going since the 1.2 days, upgraded over time, and I haven't done a lot of poking around with internals in years. My upgrade process is usually upgrading the Debian packages, merging configs with meld, running install -dis, and moving on with my day. The upgrade to Horizon 21.0 gave me a little conundrum that I was unable to resolve over the weekend. The new landing page has a status display that insisted that a variety of my nodes had outages, showing different information than 'Nodes with Outages.' Following the links to the nodes listed as down, the individual node pages directly told me that the nodes themselves were up and had 100% availability for the last 24 hours. So I decided that the status display was a 'dot zero,' probably had some issues, and simply removed it in opennms.preferences. I kept playing around and noticed that the regional status map has the same issue. The default alarm view is accurate, but if I toggle "calculate status based on outages" I see the same erroneous outage information with the same nodes as the new graphs. Nodes reported down that haven't been down in quite some time. These nodes aren't shown as being down in the outages page (/opennms/outage/list.htm?outtype=current). I did a query on the outages table, and the nodes in question are not marked as down. The outages table matches the "nodes with outages" pane on my display. (And I do currently have some outages, so I have live data to look at.) If the regional status map and the new graphs are finding outages that aren't in the outage table, they must use some other mechanism to determine that a node is down. I'm guessing that I have a table or tables, other than the outages table, that are out of sync and need to be manually fixed. Google tells me that these might be alarms that weren't properly cleared or acknowledged, or might be events that are still open. I can't see them via the GUI so I assume I have to go in with Postgresql and fix them manually. Where should I start looking? Thanks! James Zuelow OpenNMS Web Console Version: 21.0.0 Server Time: Mon Oct 30 08:09:06 AKDT 2017 Client Time: Mon Oct 30 2017 08:09:06 GMT-0800 (Alaskan Standard Time) Java Version: 1.8.0_141 (Oracle Corporation) Java Runtime: OpenJDK Runtime Environment (1.8.0_141-8u141-b15-1~deb9u1-b15) Java Specification: Java Platform API Specification (Oracle Corporation, 1.8) Java Virtual Machine: OpenJDK 64-Bit Server VM (Oracle Corporation, 25.141-b15) Java Virtual Machine Specification: Java Virtual Machine Specification (Oracle Corporation, 1.8) Operating System: Linux 4.9.0-4-amd64 (amd64) Servlet Container: jetty/9.4.0.v20161208 (Servlet Spec 3.1) User Agent: Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:56.0) Gecko/20100101 Firefox/56.0 Database Type: PostgreSQL Database Version: 9.6.4 Time-Series Strategy: RRDTool or JRobin (Interesting that my Windows workstation thinks it is standard time, not daylight savings, although it has the correct GMT offset for daylight savings.) ------------------------------------------------------------------------------ Check out the vibrant tech community on one of the world's most engaging tech sites, Slashdot.org! http://sdm.link/slashdot _______________________________________________ Please read the OpenNMS Mailing List FAQ: http://www.opennms.org/index.php/Mailing_List_FAQ opennms-discuss mailing list To *unsubscribe* or change your subscription options, see the bottom of this page: https://lists.sourceforge.net/lists/listinfo/opennms-discuss