RE: Fast multi-task abort model

"Ken Sandars" <[email protected]>
Newsgroups gmane.ietf.ips
Message-ID <[email protected]>
Hi Bill,

I'm thinking a sequence reception timeout for the outstanding TTT's will
come into effect for your scenario. However, there's always the fun of
determining how long that timeout needs to be...

Also, the target might use Nop-In ping PDUs to detect path loss - dead
initiators send no responses. This would lead to connection transport
failure, another timeout and then more clearing effects. Again, how quickly
this happens with respect to boiling HCT test scenarios is another matter.

HTH,
Ken

> -----Original Message-----
> From: [email protected] [mailto:[email protected]] On 
> Behalf Of William Studenmund
> Sent: 06 December 2005 16:17
> To: [email protected]
> Cc: [email protected]
> Subject: Re: [Ips] Fast multi-task abort model
> 
> 
> On Dec 3, 2005, at 11:06 AM, [email protected] wrote:
> 
> >>> I would let David comment on the RFC revision part -
> >>> but given all the other things we have been collecting
> >>> in the iSCSI implementer's guide and given the guide
> >>> will be published as an RFC that "updates" RFC 3720, I
> >>> think it should be OK to include the new behavior in
> >>> the guide.
> >>
> >> Just to be clear, I like both the new abort proposal 
> (AsyncEvent 5) 
> >> and
> >> also removing all long-wait inter-nexus delays. Obviously the
> >> abort/reset will have to have started on all nodes before 
> the TMF is
> >> returned. :-)
> >
> > Yes, as long as this is fully backwards compatible (which will
> > be achieved via negotiation of the next text key on login), it's
> > fine to include in the implementer's guide.
> 
> My main concern is that there is a DoS issue lurking in the 
> old method 
> we have in RFC 3720. I think there is a change we can make, to remove 
> inter-nexus delays waiting on the _completion_ of abort 
> handling, that 
> will not really be visible to the initiators yet will remove the DoS.
> 
> Here's my thought process:
> 
> I'm assuming that whatever we do with Clear Task Set/LU Reset/Target 
> Warm Reset to clear tasks is the exact same thing we will do to tasks 
> in face of a Persistent Reserve Out Preempt and Abort operation.
> 
> So here's the scenario. We have a SAN with a primary initiator and a 
> failover initiator for a specific application/service. To ensure 
> exclusive access to the target, the primary initiator is using some 
> form of reservation to keep others out of the LU(s), either 
> RESERVE/RELEASE or PERSISTENT RESERVE OUT/IN.
> 
> Now say the primary server fails and thus the failover 
> initiator has to 
> take over.
> 
> One of the first things it has to do is get the reservation. And it 
> will also want to clean up any half-completed tasks.
> 
> If we have a RESERVE/RELEASE reservation, the failover 
> initiator has to 
> issue a TMF Target Warm Reset to clear the reservation. Note: the 
> Windows HQL logo program for iSCSI requires this, so this reservation 
> system isn't dead. :-)
> 
> If we had a PERSISTENT RESERVE reservation, I believe we will want a 
> PREEMPT AND ABORT service action.
> 
> Either way, we will end up aborting all the old tasks. And we 
> have the 
> recover/failover initiator triggering a third-party cleanup on the 
> session to the failed initiator.
> 
> With the inter-nexus delay, the recovery initiator can't pick up for 
> the failed initiator until the failed initiator completes its 
> outstanding TTTs.
> 
> The failed initiator, by definition, won't be cleaning up its TTTs. 
> Thus nothing ever happens, thus a Denial of Service.
> 
> So how in the world is this supposed to work?
> 
> Am I missing something? I'd love it if someone could point it out!
> 
> I understand that there are other ways you can make a 
> cluster. However 
> I think the two cases described above are things that we should 
> support. But there's no way that can happen with what we have now.
> 
> 
> Given the above, I see no way that we can make the 
> inter-nexus delay in 
> RFC 3720 work as-is.
> 
> Thus I believe we should unilaterally change it. I think it is 
> sufficient to change it so that the TMF only has to wait for 
> termination processing to have _started_ on other sessions. In the 
> example above, if the recovery/failover initiator only had to 
> wait for 
> the failed initiator's TTTs to be marked to-die, then the TMF Target 
> Warm Reset or the Persistent Reserve Out Preempt and Abort service 
> action will proceed quickly, the tasks will be dead, and the failover 
> can proceed.
> 
> Also, while this is a notable change for targets, I do not 
> believe that 
> initiators will really notice. Can anyone else see a way that they 
> will? Well, that it could cause hardship?
> 
> > Taking off my WG Chair hat, I think the new abort proposal 
> (AsyncEvent 
> > 5)
> > is a good idea, but I'm going to be looking to Mallikarjun, Bill, 
> > Julian,
> > and others on the list to work through all the details.
> 
> I think it's a good idea too, and agree we need to look at details.
> 
> Take care,
> 
> Bill
> 
> 



_______________________________________________
Ips mailing list
[email protected]
https://www1.ietf.org/mailman/listinfo/ips
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.