Re: Silent Failure -- Enhancement Request

Carsten Finis <[email protected]>
Newsgroups gmane.comp.time.chrony.user
Message-ID <MW5PR20MB4521EA13EF0C508C18264DDEBE0D2@MW5PR20MB4521.namprd20.prod.outlook.com>
> (..just in case anybody else using Prometheus is reading this :)

Not related to the original request, but related to Prometheus monitoring of time synchronization in general:
We took a low level approach that catches a very broad range of problems by monitoring the node_timex_frequency_adjustment_ratio from the standard Prometheus node exporter. If there is no change to this value for a certain period (90 minutes in our case), there is nobody governing the kernel clock. This not only catches everything from chrony being down to servers unavailable, but it also works out of the box for other time synchronization systems like plain old ntpd.
When this alert goes off, it's usually pretty straight-forward to find the root case.

Regards,
  Carsten
-- 
To unsubscribe email chrony-users-request-kWFZVVI9zxvPqho9SqqRMmD2FQJk+8+b@public.gmane.org 
with "unsubscribe" in the subject.
For help email chrony-users-request-kWFZVVI9zxvPqho9SqqRMmD2FQJk+8+b@public.gmane.org 
with "help" in the subject.
Trouble?  Email [email protected]
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.