Re: SMP on IBM eseriesand amd64

Lukas Macura <[email protected]> Mon, 12 Jun 2006 18:13:02 +0200
Newsgroups gmane.os.openbsd.smp
Message-ID <[email protected]>
Thanks to all who are helping me!

Yes, it is strange. I heard about this that this machine should have bigger
throughput. But it does not.. Maybe we will find some bottleneck..

I tried to put some data across the box. Box is slowing connection because
server was able to achieve 40Mbytes/sec traffic if it did not go across our
box.
We have three interfaces (bge0 ad bge1 are input and ouptut ifaces of 
firewall,
on em0, there are vlans (maybe vlans are slowing communication??) and there is
DMZ interface vlan4. You see that maximum throughput is around 13Mbytes/sec.

ifstat -i bge0,bge1,em0
       bge0                bge1                em0
KB/s in  KB/s out   KB/s in  KB/s out   KB/s in  KB/s out
1140.61   1482.30   1846.41  13672.80  12701.09    677.83
1961.63   1906.47   2266.95  14208.18  13021.33   1301.74
1544.95   1913.62   2515.52  14022.09  12926.79   1226.91
1084.10   1814.52   2377.34  13460.40  12318.63    669.62
1102.96   1838.74   2237.76  13447.68  12423.09    637.19
1013.89   1772.47   2314.35  13391.93  12300.46    626.47
1084.60   1593.75   2305.87  13501.58  12369.76    828.49
1066.76   1725.85   2146.51  13557.00  12447.68    547.13
  998.55   1577.78   2157.61  13476.19  12398.57    666.44
1125.56   1702.79   2353.74  13546.82  12380.44    774.68
1349.65   1671.96   2164.74  13813.20  12460.98    667.28
1155.46   1969.64   2501.87  13714.57  12567.29    716.43
  979.44   2056.51   2591.37  13261.97  12266.25    680.06
1126.17   2064.37   2437.21  13480.67  12362.19    533.83
1229.66   2086.62   2715.48  13454.45  12257.00    833.42
1166.42   1984.17   2503.31  13574.11  12516.47    802.70
1141.74   1820.94   2279.01  13584.51  12443.09    632.97
1064.19   1972.24   2494.15  13281.63  12261.40    739.20
1724.37   1999.30   2470.15  14284.73  13271.40   1364.62
1061.01   2081.08   2596.49  13434.53  12417.23    733.66
  984.57   1849.00   2180.34  13182.10  12117.27    405.91
  982.43   2310.92   2583.39  13032.67  12008.25    387.75
1157.07   2538.57   3016.93  13463.59  12438.74    782.56
1338.31   2289.07   2530.70  13566.12  12536.73    715.35
1134.96   2159.54   2464.83  13434.33  12355.31    529.81
1356.44   1943.62   2311.49  13513.62  12323.27    692.58
1024.38   1819.81   2285.28  13275.06  12288.39    666.19
1208.24   1757.76   2261.24  13425.95  12333.64    767.69
  966.42   1702.23   1965.50  13444.56  12452.02    380.01
1012.70   1780.69   2173.39  13175.06  12131.25    514.55
  998.22   1819.66   2416.99  13358.15  12309.87    707.30
1029.78   1939.76   2459.17  13429.32  12416.67    706.06
  908.04   1974.43   2608.94  12978.15  12042.42    771.12
  894.58   1935.90   2485.24  13197.27  12287.65    705.77
  848.73   1757.40   2192.18  12844.69  11978.21    572.72
  600.02   1878.60   2108.00  13096.39  12505.82    378.97
  987.49   1648.36   2361.47  13334.47  12350.65    882.62
  980.52   1847.26   2259.06  13469.34  12554.95    646.34
  912.79   1796.95   2122.29  13211.95  12258.95    434.39
  658.41   1648.06   1847.83  12798.11  12122.58    315.53

I do not know what is low quality adapter.. These are my adapters:

em0 at pci4 dev 1 function 0 "Intel PRO/1000MT (82545GM)" rev 0x04: 
apic 12 int
0 (irq 10), address 00:0e:0c:9c:07:13

bge0 at pci5 dev 0 function 0 "Broadcom BCM5721" rev 0x11, BCM5750 B1 
(0x4101):
apic 14 int 16 (irq 10) address 00:14:5e:0b:3e:ea

bge1 at pci6 dev 0 function 0 "Broadcom BCM5721" rev 0x11, BCM5750 B1 
(0x4101):
apic 14 int 16 (irq 10) address 00:14:5e:0b:3e:eb

vmstat 1
procs   memory        page                    disks     traps         cpu
r b w    avm    fre   flt  re  pi  po  fr  sr sd0 cd0  int   sys   cs us sy id
1 3 0  98944 719784    69   0   0   0   0   0   8   0 3152  1461  189  0 11 89
1 3 0  98944 719784    16   0   0   0   0   0   0   0 13199   952  103  
0 22 78
0 3 0  98944 719784    11   0   0   0   0   0   0   0 13332  1010  127  
0 31 69
0 3 0  98944 719784     7   0   0   0   0   0   0   0 13294   868  102  
0 31 69
0 3 0  98944 719784     7   0   0   0   0   0   0   0 13420  1176  143  
0 27 73
0 3 0  98944 719784     7   0   0   0   0   0   0   0 13390  1073  136  
0 31 69
0 3 0  98944 719784     7   0   0   0   0   0   0   0 13409  1050   98  
0 27 73
0 3 0  98944 719784    11   0   0   0   0   0   0   0 13277  1064  124  
0 24 76
0 3 0  98952 719772     9   0   0   0   0   0   0   0 13359  1084  111  
0 29 71

systat vmstat shows arround 8k interrupts per iface

pfctl -vvvs i
Status: Enabled for 4 days 22:13:53           Debug: Urgent

Hostid:   0xbd4a2e30
Checksum: 0x9c14bb48b1ac30341e7da19edc7c32d3

Interface Stats for bge0              IPv4             IPv6
  Bytes In                    345357698409          8546928
  Bytes Out                   419614550986          1373749
  Packets In
    Passed                       538779628            82385
    Blocked                        5625101                0
  Packets Out
    Passed                       589144477            14869
    Blocked                         671466                0

State Table                          Total             Rate
  current entries                    10477
  searches                      3232606557         7594.8/s
  inserts                         21975946           51.6/s
  removals                        21965469           51.6/s
Source Tracking Table
  current entries                        0
  searches                               0            0.0/s
  inserts                                0            0.0/s
  removals                               0            0.0/s
Counters
  match                         1307766710         3072.5/s
  bad-offset                             0            0.0/s
  fragment                            1030            0.0/s
  short                                349            0.0/s
  normalize                              0            0.0/s
  memory                                 0            0.0/s
  bad-timestamp                          0            0.0/s
  congestion                         25715            0.1/s
  ip-option                           4593            0.0/s
  proto-cksum                        23052            0.1/s
  state-mismatch                    534903            1.3/s
  state-insert                         199            0.0/s
  state-limit                            0            0.0/s
  src-limit                              0            0.0/s
  synproxy                         1718874            4.0/s
Limit Counters
  max states per rule                    0            0.0/s
  max-src-states                         0            0.0/s
  max-src-nodes                          0            0.0/s
  max-src-conn                           0            0.0/s
  max-src-conn-rate                      0            0.0/s
  overload table insertion               0            0.0/s
  overload flush states                  0            0.0/s

Congestion of internet link is not problem, I know this. Our line is dedicated
1Gpbs full duplex. We are not using hubs :) but cisco 6550 ;) It has little
more inteligence than HUB. I hope :))



sysctl net.inet.ip.ifq.maxlen
net.inet.ip.ifq.maxlen=250
This was already set before, I found it on some conferrence.

Thank you for any next suggestions!

Lukas Macura
UIT


Quoting "Berk D. Demir" <[email protected]>:

> Lukas Macura wrote:
>> Thank you to your answer, now I know that it has no sense to compile
>> some new kernels and spend my time :)
>>
>> Our machine is used as firewall, so we really need to pin irqs to both
>> cpu and to utilise both CPUs. In this situation, we do not achieve 
>> even 100mbit throughput :( Do you
>> think it is normal? Is there any other optimalization ? Is ther
>> possibility to use first cpu for kernel and interrupts and second for
>> applications? Now only one cpu is utilized..
>
> 100mbit is an easy goal to achive with even small PCs.
>
> If you use MP kernel, APIC support is enabled and interrupt load on 
> the CPU decreases dramatically. Using MP kernel even on uniprocessor 
> systems offloads the interrupt load.
>
> If you're not using a very low quality ethernet adapter, it'll be 
> very much possible to handle ~400mbit/s traffic or ~60K pkts/s loads 
> with uniprocessor but APIC enabled workstations.
>
> Use the command "systat vmstat" and watch the interrupts columns on 
> the right. It'll display your network adapter's interrupts. Look for 
> it's iface name. (em0, sk0, bge0, etc.)
>
> Normalization, modulation and other features of PF can create 
> significant CPU load. To get detailed info about PF status, use the 
> command "pfctl -vvvs i"
>
> "congestion" is your enemy. If it keeps going up, you're probably 
> using a crappy network adapter or your switch (you're not using hubs 
> eh?) is malfunctioning.
>
> To avoid congestion you can bump "net.inet.ip.ifq.maxlen" to a value 
> your interface card can handle. I'm using value of 250 for em(4) 
> cards which are generally "Intel PRO/1000MT Dual Port Server Adapter 
> (PWLA8492MT)"
>
> Blindly bumping the number won't help but worsen the situation. I'm 
> not an expert on network adapter specs. so you have to search it for 
> yourself.
>
> Anyway, if you could reproduce the hogged scenario and post here the 
> outputs of recently mentioned commands, maybe we can find a clue.
>
> Hope this helps,
> bdd
>
>



----------------------------------------------------------------
This message was sent using IMP, the Internet Messaging Program.