Re: sys/cnmac: fix command buffer leak on transmit failure
Kirill A. Korinsky <[email protected]>
| Newsgroups | gmane.os.openbsd.tech |
|---|---|
| Message-ID | <[email protected]> |
On Fri, 10 Jul 2026 13:18:13 +0200, Kirill A. Korinsky <[email protected]> wrote: > > On Thu, 09 Jul 2026 17:55:03 +0200, > Kirill A. Korinsky <[email protected]> wrote: > > > > On Thu, 09 Jul 2026 16:22:51 +0200, > > Visa Hankala <[email protected]> wrote: > > > > > > On Thu, Jul 09, 2026 at 12:28:47AM +0200, Kirill A. Korinsky wrote: > > > > visa@, > > > > > > > > my ER-4 whcih I used as router had reached interesting state: it floods > > > > errors like this: > > > > > > > > Jun 30 11:11:39 gw /bsd: cnmac2: cannot allocate command buffer from free pool allocator > > > > Jun 30 11:11:39 gw /bsd: cnmac2: failed to transmit packet > > > > Jun 30 11:11:40 gw /bsd: cnmac2: cannot allocate command buffer from free pool allocator > > > > Jun 30 11:11:40 gw /bsd: cnmac2: failed to transmit packet > > > > > > > > tons of them, I have no idea what had happened before I discovered it in > > > > that state, but device was ok via serial, but network was dead. > > > > > > > > After reading code near that error I think I had spotted a leak. > > > > > > It seems to me that there is no leak in the code. The local cmdptr > > > pointer is derived using cmdptr_idx and discarded when the function > > > returns. cnmac_send_cmd() updates cmdptr_idx only if it was able to > > > add the command words to the queue. > > > > > > > Do you have any idea why it had exploded? > > > > What I had found in logs so far. > > That device uses 3 cnamc: > - cnmac1 is heavy used to route vlans traffic > - cnmac2 is wan port connected directly to modem > - cnmac3 is backup wan connected to the same switch to which cnmac1 is > connected > > Usually no traffic via cnmac3 until ifstated tries to switch to use it as > main wan interface. > > I don't have whole logs, just between 7 am and about 11 am where I rebooted > it, but that log states that "cannot allocate command buffer from free pool > allocator" had following distribution: > - cnmac1 -- 0 records; > - cnmac2 -- 3108 records; > - cnmac3 -- 14590 records. > > Next, I do have logs from switch but only for cnmac3 connected port (have no > idea why), and it states: > - Jun 30 at 12:25 AM device was disconnected; > - Jun 30 at 11:17 AM device was connected again. > > Unfortently I haven't tried to disconnect port physically and I do not > rememeber does a led lighted up. > > So far it smells that something happened and it consumed all PKO queue. > > But I do not understand, yet, how can it be and how can it lead to that > state which I saw. > I have added kn@ to the loop because I recall what he had reported in bugs@ the same issue with his device. I think fix is increase OCTEON_POOL_NELEMS_CMD OpenBSD currently provides 32 command buffers, each containing 127 usable 64 bit words. A packet command occupies two words, while the last word links the next buffer; one buffer therefore accommodates 63 packet commands. Every configured PKO queue owns a current tail buffer. Each interface permits 1023 unfinished transmissions or uses 17 command buffers. So, two "stopped" interface queue requires 34 buffers, what is larger than 32 as a global limit. My device with 4 interfaces seems need 68 buffers as safe limit. kn@'s with 6 102. Probably we need to bump that OCTEON_POOL_NELEMS_CMD to 128 or 256. The last one is safe for 16 interfaces. visa@, am I wrong? -- wbr, Kirill