Re: [IPFIX] draft-ietf-ipfix-a9n
Rahul Patel <[email protected]>
| Newsgroups | gmane.ietf.ipfix |
|---|---|
| Message-ID | <CB692B9A.174EF%[email protected]> |
Hi Brain, "This draft already makes a forward reference to a future correlation draft. Given that a correlation draft would at the earliest be ready for publication in the late 2013-2014 timeframe, it is IMO unrealistic to hold the aggregation draft up waiting for it by referencing it normatively." [RP] agree. Regarding Direct distribution - I am not sure if it covers the synchronized method. The description in direct distribution is not enough. May be an example would help. Regarding the late counter/active timeout - In most cases shorter active timeout would work. Thanks -Rahul On 2/21/12 10:55 AM, "Brian Trammell" <[email protected]> wrote: >Hi, Rahul, > >Comments on comments on comments... :) > >On Feb 21, 2012, at 4:34 PM, Rahul Patel wrote: > >> Hi Brian, >> >> Comments inline..[RP] >> >> On 2/21/12 5:23 AM, "Brian Trammell" <[email protected]> wrote: >> >>> Hi, Rahul, >>> >>> Thank you very much for the comments! Comments thereon inline. >>> >>> On Feb 20, 2012, at 9:07 PM, Rahul Patel wrote: >>> >>>> Hi, >>>> >>>> Here are my comments. >>>> >>>> I think correlation is important part of aggregation and this draft >>>> should >>>> include a section on it. >>>> There are different cases of correlation which the draft should touch. >>> >>> I'm not sure I agree, but it depends on what you mean by correlation. >>> More details below. >>> >>>> 1. Key fields of Aggregated Flow is same or the subset of all original >>>> Flows >>>> * In this case the correlation is straight forward. The values from >>>> the original flows are simply merged (added/replaced/etc) into >>>>matching >>>> aggregated flow. >>> >>> There is a _lot_ of complexity hiding in that "added/replaced/etc". For >>> "added" (that is, the reduced flow key as per section 5.2 of -a9n >>> represents a different Original Flow, i.e., a completely different set >>>of >>> observed packets from the point of view of the sockets at the >>>endpoints) >>> it is so straightforward it's already covered: this isn't really >>> correlation, but spatial aggregation, in the terminology of this >>>document. >>> >>> For "replaced" this requires the correlation process to know the >>> configuration of the Original Exporters so that it can differentiate >>> Original Flows representing different observed packets from the >>>endpoint >>> POV from Original Flows representing the same packets. This is where >>> correlation gets interesting, and it's my opinion that it's also out of >>> scope for this document. >>> >>> For "etc" -- here's where it gets interesting, because there might be >>> significant analysis behind trying to figure out what you should do >>>with >>> a set/cluster of N flows that look like each other. And now we're out >>>of >>> scope for the WG, IMO. >> >> [RP] 'added/replaced/etc' referred to non-key fields operation during >> aggregation. Counters will be added, min and max values may be replaced, >> averages would be recalculated, etc. >> >>> >>>> 2. Key fields of Aggregated Flow is larger than at least one original >>>> flow. >>>> * An example is best to describe this case. One original flow >>>> contain >>>> (5-tuple, src-interface, byte) and another original flow contains >>>> (5-tuple, transaction-delay). Now the aggregated flow has the key >>>>fields >>>> of srcip, dstip and src-interface). One of the flow doesn't have >>>> information on per source-interface basis. So the aggregator will >>>> perform >>>> aggregation at two level and represent that data in ipfix-structured >>>> format. The above data would look like >>>> >>>> Srcip, dstip, transaction_delay, { List of {src-interface, byte } } >>>> ===== ===== ================= ================================== >>>> A, B, 200, { (Ethernet0/0, 1500), >>>> (Ethernet1/0, 1600) .. } >>>> A, C, 300, { (Ethernet0/0, 1500) } >>> >>> Correlating {5-tuple, src-interface, octet-count} and {5-tuple, delay} >>> into {5-tuple, src-interface, octet-count, delay} is a completely >>> different operation than aggregating {5-tuple, src-interface, >>> octet-count, delay} into {src address, delay}, for example, and it >>>really >>> doesn't make any sense to me to try to wedge these into the same >>> framework. Sure, it makes sense for a lot of applications to correlate >>> before aggregating (and the document touches on this already) but that >>> doesn't mean they're the same operation. >>> >>> I will note that this second example (tweaked; see next para) would be >>>a >>> good basis for a new draft on an intermediate correlation process. But >>>I >>> don't think that draft is this draft. >>> >>> (As an aside, I find this example scary because it hides a key flow >>> attribute, byte count, behind 6313 structured data, without any >>> particular reason to do so. I am concerned that such overuse of 6313 >>> poses a serious threat to the interoperability and adoption of IPFIX. >>>But >>> that is out of scope for this email.) >>> >>> So, with respect to the inclusion of these comments in the draft, I am >>> strongly disinclined to do so. As author, I'll note that we decided >>>that >>> correlation was explicitly out of scope for this draft (paragraph 4 of >>> section 3). As a WG contributor, I'll note that correlation is >>>explicitly >>> not mentioned in the charter item for this draft, and that correlation >>> and aggregation are treated separately in 5982 and 6183. So repurposing >>> this draft at this point would go against what I believe is the sense >>>of >>> the WG to date with respect to the relationship between correlation and >>> aggregation, so I wouldn't be comfortable creeping the scope of the >>>draft >>> without a recharter. >> >> [RP] I am fine with a different draft and reference it in this draft. > >This draft already makes a forward reference to a future correlation >draft. Given that a correlation draft would at the earliest be ready for >publication in the late 2013-2014 timeframe, it is IMO unrealistic to >hold the aggregation draft up waiting for it by referencing it >normatively. > >>>> Section 5.1.1 Distributing values across Interval >>>> ------------------------------------------------- >>>> * should add 'synchronized expiry of Original Flows' - Under this >>>>method >>>> Flow C would be exported three times with start-time/end-time matching >>>> the >>>> aggregation interval start-time/end-time. >>> >>> I don't see how the result of that flow is an Aggregated Flow as per >>> section 2. >> >> [RP] It is an aggregated flow. It is very similar to Proportional >>Uniform >> Distribution (section 5.1.1) but it is more accurate >> >> | | | | >> | |<--Flow A-->| | | | >> | 10 | | | >> | | | | >> | |<--Flow B-->| | | >> | 6 | 2 | | <== Bytes in >>each >> interval >> | | | | >> | |<-------------Flow C-------------->| | >> | 4 | 32 | 10 | <== Bytes in >>each >> interval >> | | | | >> | interval 0 | interval 1 | interval 2 | >> >> >> >> Using Synchronized method >> ========================= >> Aggregated Flow will have following bytes >> Interval 0 = (10 + 6 + 4) = 20 >> Interval 1 = (0 + 2 + 32) = 34 >> Interval 2 = (0 + 0 + 10) = 10 >> >> >> Using Proportional Uniform Distribution, >> ======================================== >> (assuming Interval is 30 seconds >> Flow B overlapped 15 secs of interval 0, 7.5 secs interval 1) >> Flow C overlapped 10 secs of interval 0, 30 secs of interval 1, 20 secs >> interval 2) >> >> Aggregate Flow will have following bytes >> Interval 0 = (10 + (6+2)*(15/22.5) + 46*(10/60)) = 10+5.33+7.66 = 23 >> Interval 1 = (0 + (6+2)*(7.5/22.5) + 46*(30/60)) = 0+2.66+23 = 25.66 >> >> Interval 0 = (0 + 0 + 46*(20/60)) = 0+0+15.33 = 15.33 > >Ah. Okay. I see what you mean. Since in this case the Aggregation Process >has access to the original packet timings from the packets making up the >Original Flow, is this not covered by Direct distribution? > >>>> * There will always going to some original flows which will arrive >>>>late >>>> and will not be accounted. It will be good have a cumulative counter >>>>of >>>> such late arrivals on per template id (aggregated flow) basis. This >>>> cumulative count should be periodically exported along with end-time >>>> (time-of-snapshot). The counter will provide a degree of confidence in >>>> the >>>> data exported by the aggregator for each template. >>> >>> You can always keep this from happening by lengthening the time you are >>> willing to wait for aggregation for active + idle + export queue delay >>>+ >>> 0.5RTT + collector queue delay; in most sane cases the active timeout >>> should dominate this. >>> >>> It probably makes sense to note this in section 6 though. >> >> [RP] Some device may not want hold the data for longer to avoid larger >> memory foot print. In those case it may be desired to have such >>counters. >> I guess one could say that it is implementation detail. > >Yes, but not holding the data for longer is basically equivalent to >having a shorter active timeout, no? > >Best regards, > >Brian _______________________________________________ IPFIX mailing list [email protected] https://www.ietf.org/mailman/listinfo/ipfix