https://toitlang.org/ logo
Regular DEADLINE_EXCEEDED when using MQTT
# help
p
Is anyone else getting a lot of DEADLINE_EXCEEDED in conjunction with mqtt using cellular, like this: [cellular] DEBUG: <- +QIRD [1320] [upload.mqtt] DEBUG: closing connection {reason: DEADLINE_EXCEEDED} [cellular] DEBUG: -> AT+QICLOSE=0,0 [cellular] DEBUG: -> AT+QICLOSE=1,0 ****************************************************************************** Decoding by
jag
, device has version ****************************************************************************** EXCEPTION error. AT_COMMAND_TIMEOUT 0: Session.send. \base\at.toit:290:22 1: Session.send_ \base\at.toit:334:22 2: Session.send \base\at.toit:289:12 3: UdpSocket.close. \modules\quectel\quectel.toit:309:14 4: Locker.do. \base\at.toit:511:12
f
The stacktrace seems to be from this line (not 100% certain, since 'main' and your code seem to be different versions): https://github.com/toitware/cellular/blob/f93494904352edd71746024a07c0af5c6b1b327b/src/modules/quectel/quectel.toit#L398C28-L398C33 So basically, we don't seem to get the OK back when shutting down the modem. Not clear what that could be. Maybe earlier errors that bring the modem into a bad state. My guess is that there is a timeout when writing a message, which leads to the
[upload.mqtt] DEBUG: closing connection ...
line, which in turn tries to shut down the modem, which then too hits a timeout. So fundamentally it looks like the modem is not responding. The information here doesn't give enough information on what could be the reason. What version of the cellular package are you using? (just to know whether there were bug-fixes that have been committed since then). Also: is that something new, or did the behavior change with a recent update?
p
I'm using this branch of the cellular https://github.com/toitware/cellular/tree/kasperl-merge-upstream/src This branch have not been updated for the past 5 months
f
@kasperl Do you remember if there were important bug-fixes that were committed to main since then?
p
Hi @floitsch do we have an update on this issue?
f
Not really. @kasperl ping. Also, in order to debug this we need to know whether this is something that is new or if that behavior was always there. From the information we have here it just looks like some AT commands are timing out.
p
It has always been there. It is causing a lot of reconnects on our devices as they are not able to empty their telemetry and incident channels and is a bit annoying 🙂
f
Does it happen on every device? Does it depend on the provider? Do all logs have the same sequence of events that lead to this error? Is there anything that sticks out that could cause this issue? Fundamentally we can't really do anything without more information or access to a device that reproduces this. (Unless Kasper has an idea).
p
It happens on all devices regardless of the network provider. US, MY, AUS, EU same. I can come by and show it to you or @kasperl if you want?
f
We just need a device and instructions on how to repro.