Re: Second WGLC on draft-ietf-bmwg-traffic-management
"MORTON, ALFRED C (AL)" <[email protected]>
| Newsgroups | gmane.ietf.bmwg |
|---|---|
| Message-ID | <4AF73AA205019A4C8A1DDD32C034631D8549392B@NJFPSRVEXG0.research.att.com> |
BMWG -FYI this WGLC ends tomorrow.
Hi Barry and Ramki,
These are my comments on
Traffic Management Benchmarking
draft-ietf-bmwg-traffic-management-01.txt
Al
(as participant)
Note: as Document Shepherd, I see quite a few nits generated
by the nits-checker:
https://tools.ietf.org/idnits?url=https://tools.ietf.org/id/draft-ietf-bmwg-traffic-management-01.txt
High-level:
7. Security Considerations
8. IANA Considerations
9. Conclusions
Security and IANA sections are Mandatory, and cannot be blank.
I saw a few "TBD" in section 6, questions for Ramki. . .
The new procedures add considerable needed details, but also
raise a few questions.
When considering repeatability over multiple test runs, there could
be a statistical measure of deviation for the key benchmarks so that
the audience can easily assess the degree of consistency observed.
Detailed comments on a partial version of the draft are attached.
________________________________________
From: bmwg [[email protected]] On Behalf Of MORTON, ALFRED C (AL)
Sent: Wednesday, December 10, 2014 2:30 PM
To: [email protected]; [email protected]
Subject: [bmwg] Second WGLC on draft-ietf-bmwg-traffic-management
BMWG (and AQM):
A WG Last Call period for the Internet-Draft on
Traffic Management Benchmarking:
http://tools.ietf.org/html/draft-ietf-bmwg-traffic-management
will be open from 10 December 2014 through 6 January, 2015.
The first WGLC (on -00) closed October 21 2014, with substantial
comments and the authors believe they are now addressed.
This draft is continuing the BMWG Last Call Process. See
http://www1.ietf.org/mail-archive/web/bmwg/current/msg00846.html
Please read and express your opinion on whether or not this
Internet-Draft should be forwarded to the Area Directors for
publication as an Informational RFC. Send your comments
to this list or [email protected] and [email protected]
Al
bmwg co-chair
_______________________________________________
bmwg mailing list
[email protected]
https://www.ietf.org/mailman/listinfo/bmwg
_______________________________________________
bmwg mailing list
[email protected]
https://www.ietf.org/mailman/listinfo/bmwg
draft-ietf-bmwg-traffic-management-01acm.txt
(text/plain, 41.8 KB)
Hi Barry and Ramki,
These are my comments on
Traffic Management Benchmarking
draft-ietf-bmwg-traffic-management-01.txt
Al
(as participant)
Note: as Document Shepherd, I see quite a few nits generated
by the nits-checker:
https://tools.ietf.org/idnits?url=https://tools.ietf.org/id/draft-ietf-bmwg-traffic-management-01.txt
High-level:
7. Security Considerations
8. IANA Considerations
9. Conclusions
Security and IANA sections are Mandatory, and cannot be blank.
I saw a few "TBD" in section 6, questions for Ramki. . .
The new procedures add considerable needed details, but also
raise a few questions.
When considering repeatability over multiple test runs, there could
be a statistical measure of deviation for the key benchmarks so that
the audience can easily assess the degree of consistency observed.
Comments in context below, prefaced by [ACM]
...
4.1. Metrics for Stateless Traffic Tests
Stateless traffic measurements require that sequence number and
time-stamp be inserted into the payload for lost packet analysis.
Delay analysis may be achieved by insertion of timestamps directly
into the packets or timestamps stored elsewhere (packet captures).
This framework does not specify the packet format to carry sequence
number or timing information.
However, RFC 4737 [RFC4737] and RFC 4689 provide recommendations
for sequence tracking along with definitions of in-sequence and
out-of-order packets.
[ACM] since the reporting specs below don't include this detail, add:
. . . and Measurement Units for reported values.
<snipped BSA>. . .
- Lost Packets (LP): For all traffic management tests, the tester will
transmit the test packets into the DUT ingress port and the number of
packets received at the egress port will be measured. The difference
between packets transmitted into the ingress port and received at the
egress port is the number of lost packets as measured at the egress
port. These packets must have unique identifiers such that only the
test packets are measured. For cases where multiple flows are
transmitted from ingress to egress port (e.g. IP conversations), each
flow must have sequence numbers within the test packets stream.
RFC 4737 and RFC 2680 [RFC2680] describe the need to to establish the
time threshold to wait before a packet is declared as lost. packet as
lost, and this threshold MUST be reported with the results.
[ACM] add:
. . . and Measurement Units for reported values.
- Out of Sequence (OOS): in additions to the LP metric, the test
packets must be monitored for sequence and the out-of-sequence (OOS)
packets. RFC 4689 defines the general function of sequence tracking, as
well as definitions for in-sequence and out-of-order packets. Out-of-
order packets will be counted per RFC 4737 and RFC 2680.
[ACM] add:
. . . and Measurement Units for reported values.
- Packet Delay (PD): the Packet Delay metric is the difference between
the timestamp of the received egress port packets and the packets
transmitted into the ingress port and specified in RFC 2285. The
transmitting host and receiving host time must be in time sync using
NTP , GPS, etc.
[ACM]
Comment: I didn't find a definition for packet delay in 2285.
Is this the reference you intended? Also, need measurement units,
which will include some statistic calculated over all singleton delays.
"a real number of seconds (positive, zero or negative)"
where negative usually indicates some time sync problem, of course
- Packet Delay Variation (PDV): the Packet Delay Variation metric is
the variation between the timestamp of the received egress port
packets and specified in RFC 5481. Note that per RFC 5481, this PDV
is the variation of one-way delay across many packets in the traffic
flow.
[ACM] determine and add:
Measurement Units for reported values.
Comment: For PDV, I suggest a high percentile of 99% using the formula in
5481 and Measurement Units as similar to the units specified in 2.3 of rfc3393:
"a real number of seconds (positive or zero)"
noting that negative is no possible for PDV.
- Shaper Rate (SR): the Shaper Rate is only applicable to the
traffic shaping tests. The SR represents the average egress output
rate (bps) over the test interval.
ACM] determine and add:
Measurement Units for reported values.
- Shaper Burst Bytes (SBB): the Shaper Burst Bytes is only applicable
to the traffic shaping tests. A traffic shaper will emit packets in
different size "trains" (bytes back-to-back). This metric
characterizes the method by which the shaper emits traffic. Some
shapers transmit larger bursts per interval, and a burst of 1 packet
would apply to the extreme case of a shaper sending a CBR stream of
single packets.
[ACM] determine and add:
Measurement Units for reported values.
- Shaper Burst Interval(SBI): the interval is only applicable to the
traffic shaping tests and again is the time between shaper emitted
bursts.
[ACM] determine and add:
Measurement Units for reported values.
4.2. Metrics for Stateful Traffic Tests
The stateful metrics will be based on RFC 6349 [RFC 6349] TCP metrics and will
include:
- TCP Test Pattern Execution Time (TTPET): RFC 6349 defined the TCP
Transfer Time for bulk transfers, which is simply the measured time
to transfer bytes across single or concurrent TCP connections. The
TCP test patterns used in traffic management tests will include bulk
transfer and interactive applications. The interactive patterns include
instances such as HTTP business applications, database applications,
etc. The TTPET will be the measure of the time for a single execution
of a TCP Test Pattern (TTP). Average, minimum, and maximum times will
be measured or calculated.
[ACM] determine and add:
Measurement Units for reported values.
An example would be an interactive HTTP TTP session which should take
5 seconds on a GigE network with 0.5 millisecond latency. During ten (10)
executions of this TTP, the TTPET results might be: average of 6.5
seconds, minimum of 5.0 seconds, and maximum of 7.9 seconds.
- TCP Efficiency: after the execution of the TCP Test Pattern, TCP
Efficiency represents the percentage of Bytes that were not
retransmitted.
Transmitted Bytes - Retransmitted Bytes
TCP Efficiency % = --------------------------------------- X 100
Transmitted Bytes
Transmitted Bytes are the total number of TCP Bytes to be transmitted
including the original and the retransmitted Bytes. These retransmitted
bytes should be recorded from the sender's TCP/IP stack perspective,
to avoid any misinterpretation that a reordered packet is a retransmitted
packet (as may be the case with packet decode interpretation).
Constantine November 12, 2014 [Page 10]
Internet-Draft Traffic Management Benchmarking November, 2014
- Buffer Delay: represents the increase in RTT during a TCP test
versus the baseline DUT RTT (non congested, inherent latency). RTT
and the technique to measure RTT (average versus baseline) are defined
in RFC 6349. Referencing RFC 6349, the average RTT is derived from
the total of all measured RTTs during the actual test sampled at every
second divided by the test duration in seconds.
Total RTTs during transfer
Average RTT during transfer = -----------------------------
Transfer duration in seconds
Average RTT during Transfer - Baseline RTT
Buffer Delay % = ------------------------------------------ X 100
Baseline RTT
Note that even though this was not explicitly stated in RFC 6349,
retransmitted packets should not be used in RTT measurements.
Also, the test results should record the average RTT in millisecond
across the entire test duration and number of samples.
5. Tester Capabilities
The testing capabilities of the traffic management test environment
are divided into two (2) sections: stateless traffic testing and
stateful traffic testing
5.1. Stateless Test Traffic Generation
The test device must be capable of generating traffic at up to the
link speed of the DUT. The test device must be calibrated to verify
that it will not drop any packets. The test device's inherent PD and
PDV must also be calibrated and subtracted from the PD and PDV metrics.
The test device must support the encapsulation to be tested such as
IEEE 802.1Q VLAN, IEEE 802.1ad Q-in-Q, Multiprotocol Label Switching
(MPLS), etc. Also, the test device must allow control of the
classification techniques defined in RFC 4689 (i.e. IP address, DSCP,
TOS, etc classification).
The open source tool "iperf" can be used to generate stateless UDP
traffic and is discussed in Appendix A. Since iperf is a software
based tool, there will be performance limitations at higher link
speeds (e.g. GigE, 10 GigE, etc.). Careful calibration of any test
environment using iperf is important. At higher link speeds, it is
recommended to use hardware based packet test equipment.
Constantine November 12, 2014 [Page 11]
Internet-Draft Traffic Management Benchmarking November, 2014
5.1.1 Burst Hunt with Stateless Traffic
A central theme for the traffic management tests is to benchmark the
specified burst parameter of traffic management function, since burst
parameters of SLAs are specified in bytes. For testing efficiency,
it is recommended to include a burst hunt feature, which automates
the manual process of determining the maximum burst size which can
be supported by a traffic management function.
The burst hunt algorithm should start at the target burst size (maximum
burst size supported by the traffic management function) and will send
single bursts until it can determine the largest burst that can pass
without loss. If the target burst size passes, then the test is
complete. The hunt aspect occurs when the target burst size is not
achieved; the algorithm will drop down to a configured minimum burst
size and incrementally increase the burst until the maximum burst
supported by the DUT is discovered. The recommended granularity
of the incremental burst size increase is 1 KB.
Optionally for a policer function and if the burst size passes, the burst
should be increased by increments of 1 KB to verify that the policer is
truly configured properly (or enabled at all).
5.2. Stateful Test Pattern Generation
The TCP test host will have many of the same attributes as the TCP test
host defined in RFC 6349. The TCP test device may be a standard
computer or a dedicated communications test instrument. In both cases,
it must be capable of emulating both a client and a server.
For any test using stateful TCP test traffic, the Network Delay Emulator
(NDE function from the lab set-up diagram) must be used in order to
provide a meaningful BDP. As referenced in section 2, the target
traffic rate and configured RTT must be verified independently using
just the NDE for all stateful tests (to ensure the NDE can delay without
loss).
The TCP test host must be capable to generate and receive stateful TCP
test traffic at the full link speed of the DUT. As a general rule of
thumb, testing TCP Throughput at rates greater than 500 Mbps may require
high performance server hardware or dedicated hardware based test tools.
The TCP test host must allow adjusting both Send and Receive Socket
Buffer sizes. The Socket Buffers must be large enough to fill the BDP
for bulk transfer TCP test application traffic.
Measuring RTT and retransmissions per connection will generally require
a dedicated communications test instrument. In the absence of
dedicated hardware based test tools, these measurements may need to be
conducted with packet capture tools, i.e. conduct TCP Throughput
tests and analyze RTT and retransmissions in packet captures.
The TCP implementation used by the test host must be specified in the
test results (e.g. TCP New Reno,
TCP options supported, etc.).
[ACM] this would be a good place to add a reference to RFC3148
for examples of some of the TCP input parameters.
While RFC 6349 defined the means to conduct throughput tests of TCP bulk
transfers, the traffic management framework will extend TCP test
execution into interactive TCP application traffic. Examples include
email, HTTP, business applications, etc. This interactive traffic is
bi-directional and can be chatty.
[ACM]
by chatty, do you mean conversational, or highly interactive, or
something else?
The test device must not only support bulk TCP transfer application
traffic but also chatty traffic. A valid stress test SHOULD include
both traffic types. This is due to the non-uniform, bursty nature of
chatty applications versus the relatively uniform nature of bulk
transfers (the bulk transfer smoothly stabilizes to equilibrium state
under lossless conditions).
. . .
Application modeling techniques have been proposed in
"3GPP2 C.R1002-0 v1.0" and provides examples to model the behavior of
HTTP, FTP, and WAP applications at the TCP layer. The models have
been defined with various mathematical distributions for the
Request/Response bytes and inter-request gap times.
[ACM]
Does anyone have experience with this 3GPP modeling spec?
This framework does not specify a fixed set of TCP test patterns, but
does provide recommended test cases in Appendix B. Some of these
examples reflect those specified in "draft-ietf-bmwg-ca-bench-meth-04"
which suggests traffic mixes for a variety of representative
application profiles. Other examples are simply well-known
application traffic types such as HTTP.
6. Traffic Benchmarking Methodology
The traffic benchmarking methodology uses the test set-up from
section 2 and metrics defined in section 4.
Each test should compare the network device's internal statistics
(available via command line management interface, SNMP, etc.) to the
measured metrics defined in section 4. This evaluates the accuracy
of the internal traffic management counters under individual test
conditions and capacity test conditions that are defined in each
subsection.
From a device configuration standpoint, scheduling and shaping
functionality can be applied to logical ports such Link Aggregation
(LAG). This would result in the same scheduling and shaping
configuration applied to all the member physical ports. The focus of
this draft is only on tests at a physical port level.
The following sections provide the objective, procedure, metrics, and
reporting format for each test. For all test steps, the following
global parameters must be specified:
Test Runs (Tr). Defines the number of times the test needs to be run
to ensure accurate and repeatable results. The recommended value is 3.
Test Duration (Td). Defines the duration of a test iteration, expressed
in seconds. The recommended value it 60 seconds.
Constantine November 12, 2014 [Page 14]
Internet-Draft Traffic Management Benchmarking November, 2014
6.1. Policing Tests
Policer is defined as the entity performing the policy function. The
intent of the policing tests is to verify the policer performance
(i.e. CIR-CBS and EIR-EBS parameters). The tests will verify that the
network device can handle the CIR with CBS and the EIR with EBS and
will use back-back packet testing concepts from RFC 2544 (but adapted
to burst size algorithms and terminology). Also MEF-14,19,37 provide
some basis for specific components of this test. The burst hunt
algorithm defined in section 5.1.1 can also be used to automate the
measurement of the CBS value.
The tests are divided into two (2) sections; individual policer
tests and then full capacity policing tests. It is important to
benchmark the basic functionality of the individual policer then
proceed into the fully rated capacity of the device. This capacity may
include the number of policing policies per device and the number of
policers simultaneously active across all ports.
6.1.1 Policer Individual Tests
Objective:
Test a policer as defined by RFC 4115 or MEF 10.2, depending upon the
equipment's specification. In addition to verifying that the policer
allows the specified CBS and EBS bursts to pass, the policer test MUST
verify that the policer will remark or drop excess, and pass traffic at
the specified CBS/EBS values.
Test Summary:
Policing tests should use stateless traffic. Stateful TCP test traffic
will generally be adversely affected by a policer in the absence of
traffic shaping. So while TCP traffic could be used, it is more
accurate to benchmark a policer with stateless traffic.
As an example for RFC 4115, consider a CBS and EBS of 64KB and CIR and
EIR of 100 Mbps on a 1GigE physical link (in color-blind mode). A
stateless traffic burst of 64KB would be sent into the policer at the
GigE rate. This equates to approximately a 0.512 millisecond burst
time (64 KB at 1 GigE). The traffic generator must space these bursts
to ensure that the aggregate throughput does not exceed the CIR. The
Ti between the bursts would equal CBS * 8 / CIR = 5.12 millisecond
in this example.
Test Metrics:
The metrics defined in section 4.1 (BSA, LP, OOS, PD, and PDV) SHALL
be measured at the egress port and recorded.
Constantine November 12, 2014 [Page 15]
Internet-Draft Traffic Management Benchmarking November, 2014
Procedure:
1. Configure the DUT policing parameters for the desired CIR/EIR and
CBS/EBS values to be tested
2. Configure the tester to generate a stateless traffic burst equal
to CBS and an interval equal to Ti (CBS in bits / CIR)
3. Compliant Traffic Step: Generate bursts of CBS + EBS traffic into
the policer ingress port and measure the metrics defined in
section 4.1 (BSA, LP. OOS, PD, and PDV) at the egress port and across
the entire Td (default 60 seconds duration)
4. Excess Traffic Test: Generate bursts of greater than CBS + EBS limit
traffic into the policer ingress port and verify that the policer
only allowed the BSA bytes to exit the egress. The excess burst MUST
be recorded and the recommended value is 1000 bytes. Additional tests
beyond the simple color-blind example might include: color-aware mode,
configurations where EIR is greater than CIR, etc.
[ACM]
this is why we need Measurement Units for reported values in sec 4. . .
Reporting Format:
The policer individual report MUST contain all results for each
CIR/EIR/CBS/EBS test run and a recommended format is as follows:
********************************************************
Test Configuration Summary: Tr, Td
DUT Configuration Summary: CIR, EIR, CBS, EBS
The results table should contain entries for each test run, (Test #1
to Test #Tr).
Compliant Traffic Test: BSA, LP, OOS, PD, and PDV
Excess Traffic Test: BSA
********************************************************
Constantine November 12, 2014 [Page 16]
Internet-Draft Traffic Management Benchmarking November, 2014
6.1.2 Policer Capacity Tests
Objective:
The intent of the capacity tests is to verify the policer performance
in a scaled environment with multiple ingress customer policers on
multiple physical ports. This test will benchmark the maximum number
of active policers as specified by the device manufacturer.
Test Summary:
The specified policing function capacity is generally expressed in
terms of the number of policers active on each individual physical
port as well as the number of unique policer rates that are utilized.
For all of the capacity tests, the benchmarking test procedure and
report format described in Section 6.1.1 for a single policer MUST
be applied to each of the physical port policers.
As an example, a Layer 2 switching device may specify that each of the
32 physical ports can be policed using a pool of policing service
policies. The device may carry a single customer's traffic on each
physical port and a single policer is instantiated per physical port.
Another possibility is that a single physical port may carry multiple
customers, in which case many customer flows would be policed
concurrently on an individual physical port (separate policers per
customer on an individual port).
Test Metrics:
The metrics defined in section 4.1 (BSA, LP, OOS, PD, and PDV) SHALL
be measured at the egress port and recorded.
The following sections provide the specific test scenarios,
procedures, and reporting formats for each policer capacity test.
6.1.2.1 Maximum Policers on Single Physical Port Test
Test Summary:
The first policer capacity test will benchmark a single physical port,
maximum policers on that physical port.
Assume multiple categories of ingress policers at rates r1, r2,...rn.
There are multiple customers on a single physical port. Each customer
could be represented by a single tagged vlan, double tagged vlan,
VPLS instance etc. Each customer is mapped to a different policer.
Each of the policers can be of rates r1, r2,..., rn.
An example configuration would be
- Y1 customers, policer rate r1
- Y2 customers, policer rate r2
- Y3 customers, policer rate r3
...
- Yn customers, policer rate rn
Constantine November 12, 2014 [Page 17]
Internet-Draft Traffic Management Benchmarking November, 2014
Some bandwidth on the physical port is dedicated for other traffic (non
customer traffic); this includes network control protocol traffic. There
is a separate policer for the other traffic. Typical deployments have 3
categories of policers; there may be some deployments with more or less
than 3 categories of ingress policers.
Test Procedure:
1. Configure the DUT policing parameters for the desired CIR/EIR and
CBS/EBS values for each policer rate (r1-rn) to be tested
2. Configure the tester to generate a stateless traffic burst equal to
CBS and an interval equal to TI (CBS in bits/CIR) for each customer
stream (Y1 - Yn). The encapsulation for each customer must also be
configured according to the service tested (VLAN, VPLS, IP mapping,
etc.).
3. Compliant Traffic Step: Generate bursts of CBS + EBS traffic into the
policer ingress port for each customer traffic stream and measure the
metrics defined in section 4.1 (BSA, LP, OOS, PD, and PDV) at the
egress port for each stream and across the entire Td (default 30
seconds duration)
4. Excess Traffic Test: Generate bursts of greater than CBS + EBS limit
traffic into the policer ingress port for each customer traffic
stream and verify that the policer only allowed the BSA bytes to exit
the egress for each stream. The excess burst MUST recorded and the
recommended value is 1000 bytes.
Constantine November 12, 2014 [Page 18]
Internet-Draft Traffic Management Benchmarking November, 2014
Reporting Format:
The policer individual report MUST contain all results for each
CIR/EIR/CBS/EBS test run, per customer traffic stream.
A recommended format is as follows:
********************************************************
Test Configuration Summary: Tr, Td
Customer traffic stream Encapsulation: Map each stream to VLAN,
VPLS, IP address
DUT Configuration Summary per Customer Traffic Stream: CIR, EIR,
CBS, EBS
The results table should contain entries for each test run, (Test #1
to Test #Tr).
Customer Stream Y1-Yn (see note), Compliant Traffic Test: BSA, LP,
OOS, PD, and PDV
Customer Stream Y1-Yn (see note), Excess Traffic Test: BSA
********************************************************
Note: For each test run, there will be a two (2) rows for each
customer stream, the compliant traffic result and the excess traffic
result.
6.1.2.2 Single Policer on All Physical Ports
Test Summary:
The second policer capacity test involves a single Policer function per
physical port with all physical ports active. In this test, there is a
single policer per physical port. The policer can have one of the rates
r1, r2,.., rn. All the physical ports in the networking device are
active.
Procedure:
The procedure is identical to 6.1.1, the configured parameters must be
reported per port and the test report must include results per
measured egress port
6.1.2.3 Maximum Policers on All Physical Ports
Finally the third policer capacity test involves a combination of the
first and second capacity test, namely maximum policers active per
physical port and all physical ports are active.
Procedure:
Uses the procedural method from 6.1.2.1 and the configured parameters
must be reported per port and the test report must include per stream
results per measured egress port.
Constantine November 12, 2014 [Page 19]
Internet-Draft Traffic Management Benchmarking November, 2014
6.2. Queue and Scheduler Tests
Queues and traffic Scheduling are closely related in that a queue's
priority dictates the manner in which the traffic scheduler
transmits packets out of the egress port.
Since device queues / buffers are generally an egress function, this
test framework will discuss testing at the egress (although the
technique can be applied to ingress side queues).
Similar to the policing tests, the tests are divided into two
sections; individual queue/scheduler function tests and then full
capacity tests.
6.2.1 Queue/Scheduler Individual Tests Overview
The various types of scheduling techniques include FIFO, Strict
Priority (SP), Weighted Fair Queueing (WFQ) along with other
variations. This test framework recommends to test at a minimum
of three techniques although it is the discretion of the tester
to benchmark other device scheduling algorithms.
6.2.1.1 Queue/Scheduler with Stateless Traffic Test
Objective:
Verify that the configured queue and scheduling technique can
handle stateless traffic bursts up to the queue depth.
Test Summary:
A network device queue is memory based unlike a policing function,
which is token or credit based. However, the same concepts from
section 6.1 can be applied to testing network device queues.
The device's network queue should be configured to the desired size
in KB (queue length, QL) and then stateless traffic should be
transmitted to test this QL.
A queue should be able to handle repetitive bursts with the
transmission gaps proportional to the bottleneck bandwidth. This
gap is referred to as the transmission interval (Ti). Ti can
be defined for the traffic bursts and is based off of the QL and
Bottleneck Bandwidth (BB) of the egress interface.
Ti = QL * 8 / BB
Note that this equation is similar to the Ti required for transmission
into a policer (QL = CBS, BB = CIR). Also note that the burst hunt
algorithm defined in section 5.1.1 can also be used to automate the
measurement of the queue value.
Constantine November 12, 2014 [Page 20]
Internet-Draft Traffic Management Benchmarking November, 2014
The stateless traffic burst shall be transmitted at the link speed
and spaced within the Ti time interval. The metrics defined in section
4.1 shall be measured at the egress port and recorded; the primary
result is to verify the BSA and that no packets are dropped.
The scheduling function must also be characterized to benchmark the
device's ability to schedule the queues according to the priority.
An example would be 2 levels of priority including SP and FIFO
queueing. Under a flow load greater the egress port speed, the
higher priority packets should be transmitted without drops (and
also maintain low latency), while the lower priority (or best
effort) queue may be dropped.
Test Metrics:
The metrics defined in section 4.1 (BSA, LP, OOS, PD, and PDV) SHALL
be measured at the egress port and recorded.
Procedure:
1. Configure the DUT queue length (QL) and scheduling technique
(FIFO, SP, etc) parameters
2. Configure the tester to generate a stateless traffic burst equal
to QL and an interval equal to Ti (QL in bits/BB)
3. Generate bursts of QL traffic into the DUT and measure the
metrics defined in section 4.1 (LP, OOS, PD, and PDV) at the egress
port and across the entire Td (default 30 seconds duration)
Report Format:
The Queue/Scheduler Stateless Traffic individual report MUST contain
all results for each QL/BB test run and a recommended format is as
follows:
********************************************************
Test Configuration Summary: Tr, Td
DUT Configuration Summary: Scheduling technique, BB and QL
The results table should contain entries for each test run as follows,
(Test #1 to Test #Tr).
- LP, OOS, PD, and PDV
********************************************************
Constantine November 12, 2014 [Page 21]
Internet-Draft Traffic Management Benchmarking November, 2014
6.2.1.2 Testing Queue/Scheduler with Stateful Traffic
Objective:
Verify that the configured queue and scheduling technique can handle
stateless traffic bursts up to the queue depth.
Test Background and Summary:
To provide a more realistic benchmark and to test queues in layer 4
devices such as firewalls, stateful traffic testing is recommended
for the queue tests. Stateful traffic tests will also utilize the
Network Delay Emulator (NDE) from the network set-up configuration in
section 2.
The BDP of the TCP test traffic must be calibrated to the QL of the
device queue. Referencing RFC 6349, the BDP is equal to:
BB * RTT / 8 (in bytes)
The NDE must be configured to an RTT value which is large enough to
allow the BDP to be greater than QL. An example test scenario is
defined below:
- Ingress link = GigE
- Egress link = 100 Mbps (BB)
- QL = 32KB
RTT(min) = QL * 8 / BB and would equal 2.56 millisecond (and the
BDP = 32KB)
In this example, one (1) TCP connection with window size / SSB of
32KB would be required to test the QL of 32KB. This Bulk Transfer
Test can be accomplished using iperf as described in Appendix A.
Two types of TCP tests must be performed: Bulk Transfer test and Micro
Burst Test Pattern as documented in Appendix B. The Bulk Transfer
Test only bursts during the TCP Slow Start (or Congestion Avoidance)
state, while the Micro Burst test emulates application layer bursting
which may occur any time during the TCP connection.
Other tests types should include: Simple Web Site, Complex Web Site,
Business Applications, Email, SMB/CIFS File Copy (which are also
documented in Appendix B).
Test Metrics:
The test results will be recorded per the stateful metrics defined in
section 4.2, primarily the TCP Test Pattern Execution Time (TTPET),
TCP Efficiency, and Buffer Delay.
Constantine November 12, 2014 [Page 22]
Internet-Draft Traffic Management Benchmarking November, 2014
Procedure:
1. Configure the DUT queue length (QL) and scheduling technique
(FIFO, SP, etc) parameters
2. Configure the tester* to generate a profile of emulated of an
application traffic mixture
- The application mixture MUST be defined in terms of percentage
of the total bandwidth to be tested
- The rate of transmission for each application within the mixture
MUST be also be configurable
* The tester MUST be capable of generating a precise TCP test
patterns for each application specified, to ensure repeatable results.
3. Generate application traffic between the ingress (client side) and
egress (server side) ports of the DUT and measure the metrics (TTPET,
TCP Efficiency, and Buffer Delay) per application stream and at the
ingress and egress port (across the entire Td, default 60 seconds
duration).
Reporting Format:
The Queue/Scheduler Stateful Traffic individual report MUST contain all
results for each traffic scheduler and QL/BB test run and a recommended
format is as follows:
********************************************************
Test Configuration Summary: Tr, Td
DUT Configuration Summary: Scheduling technique, BB and QL
Application Mixture and Intensities: this is the percent configured of
each application type
The results table should contain entries for each test run as follows,
(Test #1 to Test #Tr).
- Per Application Throughout (bps) and TTPET
- Per Application Bytes In and Bytes Out
- Per Application TCP Efficiency, and Buffer Delay
********************************************************
[ACM]
I'm unsure how to calculate the "Per Application" metrics,
it's vague because the definition of Application could span
multiple sessions for different "clients", or it could be
a metric for single client performance, or ? Think about
statistical summarization, if Per App = single client.
Also, we have to make clear that this form of Throughput is
different from RFC1242/2544 Throughput.
Constantine November 12, 2014 [Page 23]
Internet-Draft Traffic Management Benchmarking November, 2014
6.2.2 Queue / Scheduler Capacity Tests
Objective:
The intent of these capacity tests is to benchmark queue/scheduler
performance in a scaled environment with multiple queues/schedulers
active on multiple egress physical ports. This test will benchmark
the maximum number of queues and schedulers as specified by the
device manufacturer. Each priority in the system will map to a
separate queue.
Test Metrics:
The metrics defined in section 4.1 (BSA, LP, OOS, PD, and PDV) SHALL
be measured at the egress port and recorded.
The following sections provide the specific test scenarios, procedures,
and reporting formats for each queue / scheduler capacity test.
6.2.2.1 Multiple Queues / Single Port Active
For the first scheduler / queue capacity test, multiple queues per
port will be tested on a single physical port. In this case,
all the queues (typically 8) are active on a single physical port.
Traffic from multiple ingress physical ports are directed to the
same egress physical port which will cause oversubscription on the
egress physical port.
There are many types of priority schemes and combinations of
priorities that are managed by the scheduler. The following
sections specify the priority schemes that should be tested.
6.2.2.1.1 Strict Priority on Egress Port
Test Summary:
For this test, Strict Priority (SP) scheduling on the egress
physical port should be tested and the benchmarking methodology
specified in section 6.2.1.1 and 6.2.1.2 (procedure, metrics,
and reporting format) should be applied here. For a given
priority, each ingress physical port should get a fair share of
the egress physical port bandwidth.
TBD: RAMKI, do we need a concrete example?
[ACM] above
Since this is a capacity test, the configuration and report
results format from 6.2.1.1 and 6.2.1.2 MUST also include:
Configuration:
- The number of physical ingress ports active during the test
- The classication marking (DSCP, VLAN, etc.) for each physical
ingress port
- The traffic rate for stateful traffic and the traffic rate
/ mixture for stateful traffic for each physical ingress port
Report results:
- For each ingress port traffic stream, the achieved throughput
rate and metrics at the egress port
Constantine November 12, 2014 [Page 24]
Internet-Draft Traffic Management Benchmarking November, 2014
6.2.2.1.2 Strict Priority + Weighted Fair Queue (WFQ) on Egress Port
Test Summary:
For this test, Strict Priority (SP) and Weighted Fair Queue (WFQ)
should be enabled simultaneously in the scheduler but on a single
egress port. The benchmarking methodology specified in Section
6.2.1.1 and 6.2.1.2 (procedure, metrics, and reporting format)
should be applied here. Additionally, the egress port bandwidth
sharing among weighted queues should be proportional to the assigned
weights. For a given priority, each ingress physical port should get
a fair share of the egress physical port bandwidth.
TBD: RAMKI, do we need a concrete example?
[ACM] above
Since this is a capacity test, the configuration and report results
format from 6.2.1.1 and 6.2.1.2 MUST also include:
Configuration:
- The number of physical ingress ports active during the test
- The classication marking (DSCP, VLAN, etc.) for each physical
ingress port
- The traffic rate for stateful traffic and the traffic rate /
mixture for stateful traffic for each physical ingress port
Report results:
- For each ingress port traffic stream, the achieved throughput rate
and metrics at each queue of the egress port queue (both the SP
and WFQ queue).
Example:
- Egress Port SP Queue: throughput and metrics for ingress streams 1-n
- Egress Port WFQ Queue: throughput and metrics for ingress streams 1-n
6.2.2.2 Single Queue per Port / All Ports Active
Test Summary:
Traffic from multiple ingress physical ports are directed to the
same egress physical port, which will cause oversubscription on the
egress physical port. Also, the same amount of traffic is directed
to each egress physical port.
The benchmarking methodology specified in Section 6.2.1.1
and 6.2.1.2 (procedure, metrics, and reporting format) should be
applied here. Each ingress physical port should get a fair share of
the egress physical port bandwidth. Additionally, each egress
physical port should receive the same amount of traffic.
Constantine November 12, 2014 [Page 25]
Internet-Draft Traffic Management Benchmarking November, 2014
Since this is a capacity test, the configuration and report results
format from 6.2.1.1 and 6.2.1.2 MUST also include:
Configuration:
- The number of ingress ports active during the test
- The number of egress ports active during the test
- The classication marking (DSCP, VLAN, etc.) for each physical
ingress port
- The traffic rate for stateful traffic and the traffic rate /
mixture for stateful traffic for each physical ingress port
Report results:
- For each egress port, the achieved throughput rate and metrics at
the egress port queue for each ingress port stream.
Example:
- Egress Port 1: throughput and metrics for ingress streams 1-n
- Egress Port n: throughput and metrics for ingress streams 1-n
6.2.2.3 Multiple Queues per Port, All Ports Active
Traffic from multiple ingress physical ports are directed to all
queues of each egress physical port, which will cause
oversubscription on the egress physical ports. Also, the same
amount of traffic is directed to each egress physical port.
The benchmarking methodology specified in Section 6.2.1.1
and 6.2.1.2 (procedure, metrics, and reporting format) should be
applied here. For a given priority, each ingress physical port
should get a fair share of the egress physical port bandwidth.
Additionally, each egress physical port should receive the same
amount of traffic.
Since this is a capacity test, the configuration and report results
format from 6.2.1.1 and 6.2.1.2 MUST also include:
Configuration:
- The number of physical ingress ports active during the test
- The classication marking (DSCP, VLAN, etc.) for each physical
ingress port
- The traffic rate for stateful traffic and the traffic rate /
mixture for stateful traffic for each physical ingress port
Report results:
- For each egress port, the achieved throughput rate and metrics at
each egress port queue for each ingress port stream.
Example:
- Egress Port 1, SP Queue: throughput and metrics for ingress streams 1-n
- Egress Port 2, WFQ Queue: throughput and metrics for ingress streams 1-n
.
.
- Egress Port n, SP Queue: throughput and metrics for ingress streams 1-n
- Egress Port n, WFQ Queue: throughput and metrics for ingress streams 1-n
Constantine November 12, 2014 [Page 26]
Internet-Draft Traffic Management Benchmarking November, 2014
<stopped here>