Re: Connection loss (IWD HEAD with latest OWE / BSS selection patches) - brcmfmac driver
Martin Petzold <[email protected]>
| Newsgroups | dev.linux.lists.iwd |
|---|---|
| Organization | TAVLA Technology GmbH |
| Message-ID | <[email protected]> |
Dear James, Am 04.11.24 um 13:36 schrieb James Prestwood: > > On 11/3/24 3:13 PM, Martin Petzold wrote: >> Dear James, >> >> Am 25.10.24 um 17:17 schrieb James Prestwood: >>> >>>>>>>>>>> >>>>>>>>>>> >>>>>>>>>>> I open a new thread for this one: During the last weeks I >>>>>>>>>>> have seen connection losses for 30+ minutes, sometimes even >>>>>>>>>>> hours or just now even forever (IWD HEAD with v2 OWE / BSS >>>>>>>>>>> selection patches). Driver is brcmfmac (NXP 6.1.36 kernel) >>>>>>>>>>> and chip is BCM4339 (Laird LWB5). >>>>>>>>>>> >>>>>>>>>>> It happens in a) single router environment (WPA2-PSK; >>>>>>>>>>> Touchstone TG3442DE), and b) router + repeater environment >>>>>>>>>>> (WPA2 CCMP; Fritz!Box + Fritz!Repeater), and maybe also in >>>>>>>>>>> the WPA3 OWE Transition network (yesterday lost a connection >>>>>>>>>>> again). >>>>>>>>>> >>>>>>>>>> I lost now again 2 of 10 devices in the WPA3 OWE network >>>>>>>>>> (with roaming). However, now they don't disappear all after a >>>>>>>>>> shorter while. It seems to be later. >>>>>>>>>> >>>>>>>>>> I also lost one device in a Router+Repeater WPA2 (CCMP) >>>>>>>>>> network. It is confirmed here on router side, that the device >>>>>>>>>> is disconnected. Since more than a day. >>>>>>>>> >>>>>>>>> We can't do anything without logs. If you suspect its the >>>>>>>>> blacklist you can lower the blacklist time down in main.conf: >>>>>>>>> >>>>>>>>> [ >> >> I am still losing devices. Sometimes they come back again, but mostly >> do not re-connect. I have observed the following: >> >> - Connection exists for several hours until about one day, or two. >> Then gone for several hours or mostly forever. >> - For FritzBox+FritzRepeater I have seen the connection coming back >> after like a day (here connection loss was also confirmed on router >> side!) >> - For the Aruba enterprise environment the connection never came back >> (until now no AP logs - waiting for an answer) >> - After reboot the connection comes back >> - It occurs only in an environment with multiple APs with same SSID >> (i.e. roaming environment), however my single AP environments have >> all strong signal >> - Some devices with identical configuration in this environment DO >> NOT get lost, those seem to have quite strong signal (maybe they >> don't roam) >> - Other devices in the same environment work without any problems >> (Intel+NetworkManager) and the APs are Aruba enterprise grade >> - I see almost the same in the Aruba enterprise environment, but ALSO >> in a FritzBox + FritzRepeater environment >> - We had a bug in our web socket connection, causing to many IWD >> requests. However, this was fixed. And why are all the other devices >> okay? Maybe co-incidence with roaming and anything related to >> dropping and re-connecting web socket connection. >> >> Please find attached my currently available debug logs (they are a >> few days old, but I am quite sure this is the connection loss >> situation). These logs are from the FritzBox+FritzRepeater >> environment. There are no brcmfmac messages (but also no special >> debug level configured here)! >> >> I have now also disabled WiFi power saving and will deploy to the >> environment...hoping the best. >> >> Maybe you could check the logs and have an idea? > > Looks like the same thing as the last logs you sent. IWD tries to > connect (sends CMD_CONNECT to the kernel) but gets no associated > CMD_CONNECT event after that which causes IWD to wait indefinitely for > that event. This, again, appears like a driver problem because its > expected that the kernel tells userspace the result of the CMD_CONNECT > request. > > Only similarity I can see between the two sets of logs is there is a > failed connection just prior to the hang. IWD then attempts to connect > again but the 4-way handshake is never started and this results in a > failure with status 16 (group key handshake timeout). In your latest > set of logs IWD actually again tries to connect to a different BSS and > gets status 16 before trying yet again and hanging. > > This actually seems similar to an issue I encountered with ath10k > where the network interface would time out being brought up. Retrying > would succeed but the driver would be in a similar state where IWD > could authenticate/associate but no data frames (i.e. 4-way handshake) > would be passed to userspace. Only solution (until upstream fixed the > bug) was to unload/reload the driver when we detected this condition. > > If you are able to physically attach to a device currently in this > state you may be able to get more info. For example if IWD is stuck > like this try disconnecting/reconnecting with iwctl or restarting IWD > to see what happens. If you end up in the same state right away I'm > 99.9% sure the driver is the entire reason your running into this. ----- Okt 24 02:30:33 tavla iwd[386]: src/netdev.c:netdev_mlme_notify() MLME notification Disconnect(48) Okt 24 02:30:33 tavla iwd[386]: src/netdev.c:netdev_disconnect_event() Okt 24 02:30:33 tavla iwd[386]: Received Deauthentication event, reason: 8, from_ap: true ----- From where does this MLME notification Disconnect come before the AP Deauth event? And does reason 8 mean "Disassociated because sending station is leaving (or has left)" (which sounds like originally station decision) OR "AP moved the client to another access point using non-aggressive load balancing" (which then seems to be an AP decision)? Because we are talking about signals around -70. What would the event be, if the AP finds the station unacceptable because of bad signal strength (so this is signal in direction to AP). Or doesn't this exist? Thanks, Martin