Re: [friam] Re: Distributed failure models

"Karp, Alan H" <[email protected]> Mon, 6 Jul 2015 16:06:14 +0000
Newsgroups gmane.comp.capabilities.general
Message-ID <8AD823089998C849A832D86972E69CD55B8E4265@G4W3222.americas.hpqcorp.net>
Daira Hopwood wrote:

> I continue to be convinced that both models are necessary and must be used simultaneously,

An excellent summary of the issues.  To me the key point you made is that a message can only be dropped when both the sender and the (potentially dead) receiver agree that it will never be processed.  That's forbidden by the waterken model but is the essence of the E model if I understand correctly.

One thing you didn't mention is adding the ability for the application layer to configure the infrastructure layer.  I see two justifications.  The first is that only the application layer has the information needed to properly set some timeouts.  The infrastructure can use this information when setting its timeouts, e.g., use the max of the application and infrastructure timeouts.

I'm less sure of the second justification, but it seems to me that only the application is in a position to know if it is processing a message of death.  If the application on the receiving side can detect that it is processing such a message, it can tell its infrastructure not to deliver that message again.  The infrastructure on the receiving side can inform the sending side, which can tell the sending application.  The sending side of the application may then be able to take corrective action, including telling its infrastructure to cancel the message.  This communication between the application and its infrastructure violates one of the fundamental principles of the waterken approach, but it seems to meet your goal of supporting a hybrid of the two approaches.

________________________
Alan Karp
Principal Scientist
Enterprise Services, Office of the CTO
Hewlett-Packard Company
1501 Page Mill Road
Palo Alto, CA 94304
1 (650) 386-4568
http://alankarp.parseapp.com