Re: comments on draft-banks-bmwg-issu-meth-01
"Fernando Calabria (fcalabri)" <[email protected]>
| Newsgroups | gmane.ietf.bmwg |
|---|---|
| Message-ID | <[email protected]> |
Erum, thanks a lot for taking the time to provide feedback on the doc ! My comments –/ toughs inline … (in addition to Sarah's reply) From: "Erum Frahim (efrahim)" <[email protected]<mailto:[email protected]>> Date: Saturday, July 27, 2013 5:24 PM To: "[email protected]<mailto:[email protected]>" <[email protected]<mailto:[email protected]>> Subject: [bmwg] comments on draft-banks-bmwg-issu-meth-01 Hi Team, Here is some of the feedback and suggestions: Missing the 5.1 in table of Content. 5. ISSU Test Methodology............................................10 5.1 Pre-ISSU recommended verifications ……………………………………………… >> ops – good catch ! 1. Introduction <efrahim> Can we define the patching and maintenance upgrade little bit more. ISSU operation can be categorized into multiple scenario - Whole Atomic version change: In this scenario, the whole system including all the process are upgraded as well in certain cases the firmware and bios can be upgraded. - Patching: This usually refer to the one or two process upgrade. Patching usually helpful to fix a bug on the current version. - Maintenance: If there are multiple fixes are coming for many modules. >> We took ISSU as a 'holistic approach' the methodology described applies for a major software upgrade as well as a software – maintenance path , but point taken, we will add some additional text on this … <efrahim> Some Datacenter layer 3 switches are referred their Routing Processor as Supervisors (SUP). It would be nice to add both reference. Different hardware configurations may be expected to be benchmarked, but a typical configuration for a forwarding device that supports ISSU consists of at least one pair of Routing Processors (RP's) or Supervisors (SUP) >> We can add a comments making a SUP module = as an RP <efrahim> Since most modern forwarding devices, where ISSU would be applicable, do consist of redundant RP's or Sups and hardware-separated control plane and data plane functionality, this document will focus on methodologies which would be directly applicable to those platforms. <efrahim> Do we really want to restrict the ISSU to only redundant supervisors or RP. A lot of the new DC switches can easily do patching and ISSU without the redundant supervisor or RP. It would be good to include both ISSU testing methodology. The procedure if there is single SUP. >> We do not really want to restrict it, this is not a mandatory statement, at the time of writing this doc and even today an ISSU process with a single RP / SUP is 'disruptive' in nature (some of us in the industry) do not consider it an 'IN SERVICE' event, while in some specific platforms this can be achieved (for a very limited type of deployments) this just applies for a very limited set of devices and deployments .>> We will discuss this internally and within the WG and let you know where we go with this .. 3.3 Upgrade Run <efrahim> May be change to include and clear the process better. Here is the suggestion In this phase, the secondary RP load the new software first. Once the software is loaded on the secondary RP and comes to HA State, the secondary RP/SUP takes over, forcing the RP/SUP which was previously designated as primary, to adopt the standby role. At this point, the new primary RP drives the required updates to other specific components and forces warm-updates or re-initializations with the new software, as applicable. In addition, the now-standby RP will be updated with the desired software. Depend of vendor implementation, after upgrading and downgrading the SUP/RP, the line cards (LC) and other modules gets upgraded either serially or in parallel. >> I personally have some concerns here, a basic one is is that the operator / tester does not necessarily needs to know that for example the STBY RP / SUP is upgraded first, maybe even software is 'stage' on the specific LCS, this is why we are just describing a 'typical' process, but we do not want to make this a restrictive requirement as different vendors may deploy – come up with new ways to achieve this >> 5.3 Upgrade Run <efrahim> There is no reference for any T3 Time where on the section 7 Final Report, there is a mention of T3 time. It would be nice to have some reference of T Time. >> We define T3 as the time for boTH SUP/RP s to me bcd up to NSR/HA synced state .. WE can add a reference to this time at the end of 5.3 … <efrahim> There was no mention of LC upgrades after the Supervisor. Upgrade processes do include both Supervisor and Lcs. May be log that as T4. At this point, pay particular attention to any indications of control plane disruption, traffic impact or other anomalous behavior. Once the DUT has converged upon the new code and returned to normal operation note the completion time and log the duration of this step as T2. >> By the end of T2 time all forwarding plane should be resumed, and reported by TP_frames , this considers the LC upgrade time (we may need to add a reference / comment here) what some on the industry call as 'dark window' I do not see the Lcs being upgrade after or taking more than T4, you lost me on this one .. :) Once both RP/SUPs are up on new code and synchronized, upgrade the LC and other modules. Log this duration of this time as T4. 7 Final Report Add the T4 time for LC and other modules upgrade. ISSU for all the LCs and other components T4 Total ISSU Maintenance Window T5 (sum of T1+T2+T3+T4) Regards Erum >> Thanks a lot for all you comments! _______________________________________________ bmwg mailing list [email protected] https://www.ietf.org/mailman/listinfo/bmwg