https://toitlang.org/ logo
Constant Out of Memory errors - How to optimize?
# help
a
I keep getting tons of memory crashes. How can I optimize my code? Its already pretty minimal so I'm not sure what to do. Crash log example:
Copy code
Heap report @ out of memory:
  ┌───────────┬──────────┬─────────────────────────────────────────────────────┐
  │   Bytes   │  Count   │  Type                                               │
  ├───────────┼──────────┼─────────────────────────────────────────────────────┤
  │     640   │      1   │  external byte array                                │
  │    2088   │     12   │  tls/bignum                                         │
  │   53248   │     15   │  toit processes                                     │
  │     16384 │        4 │    system    0 b476094f-1061-696e-a209-3f0ad2a0162d │
  │     24576 │        6 │    current   1 68a2c931-1e01-4f20-cfbd-6a64b500a955 │
  │     12288 │        3 │    other     2 836308cc-d8d0-e73d-6648-17a1bcd6d444 │
  │   16384   │      1   │  heap metadata                                      │
  │    4096   │      1   │  spare new-space                                    │
  │    6088   │     21   │  lwip                                               │
  │    9608   │    837   │  heap overhead                                      │
  │    1944   │     29   │  event source                                       │
  │   50280   │    390   │  thread/other                                       │
  │   30984   │     25   │  thread/spawn                                       │
  │   51368   │    226   │  untagged                                           │
  │   11936   │     56   │  wifi                                               │
  └───────────┴──────────┴─────────────────────────────────────────────────────┘
  Total: 238664 bytes in 775 allocations (93%), largest free 2k, total free 19k
f
@erikcorry might have suggestions. To me it looks like there is some free memory, but it's not contiguous. It also looks like the device is in the process of starting a TLS connection. Could be the one for the Artemis update check. If you still have max-offline set to 0s consider increasing it, do that you doing need a full TLS connection all the time.
Is this happening at boot time, or later?
If you want to, you can send me your code and I could look whether memory is lost somewhere.
My email is florian@toit.io.
Also. If you don't use Bluetooth you can recover some memory by using a different envelope. (That memory is reserved even if you don't use BLE)
If you add
envelope: esp32-no-ble
to your pod specification that would use the envelope without BLE support.
a
Could be the max-offline. I do need BLE though unfortunately. I'll increase the max-offline time and report back.
I don't mind sending the code here- it'll be open sourced anyways:
Copy code
toit
import ble
import net
import http
import encoding.hex
import certificate-roots
import artemis

SCAN-DURATION   ::= Duration --s=20

main args:
  certificate-roots.install-common-trusted-roots
  adapter := ble.Adapter
  central := adapter.central

  network := net.open
  client :=  http.Client network
  headers := http.Headers

  if args[1] == null or args[1] == "":
    throw "Authorization header is required"
    exit 1

  headers.add "Authorization" args[1]
  headers.add "Device-ID" artemis.device.id.to-string

  try:
    session := client.web-socket --uri=args[0] --headers=headers // ws://192.168.0.58:8080/ws

    try:
      central.scan --duration=SCAN-DURATION: | device/ble.RemoteScannedDevice |
        session.send (hex.encode device.data.manufacturer-data)
    finally:
      session.close
  finally:
      network.close
I can't think of any reason why I would continue to get OOM errors- maybe the hex encoding? OOM Error:
Copy code
Heap report @ out of memory:
  ┌───────────┬──────────┬─────────────────────────────────────────────────────┐
  │   Bytes   │  Count   │  Type                                               │
  ├───────────┼──────────┼─────────────────────────────────────────────────────┤
  │    2176   │      2   │  external byte array                                │
  │   16936   │     23   │  tls/bignum                                         │
  │   49152   │     14   │  toit processes                                     │
  │     16384 │        4 │    system    0 b476094f-1061-696e-a209-3f0ad2a0162d │
  │     24576 │        6 │    other     1 68a2c931-1e01-4f20-cfbd-6a64b500a955 │
  │      8192 │        2 │    current   4 836308cc-d8d0-e73d-6648-17a1bcd6d444 │
  │   16384   │      1   │  heap metadata                                      │
  │    4096   │      1   │  spare new-space                                    │
  │    8048   │     25   │  lwip                                               │
  │    8136   │    657   │  heap overhead                                      │
  │    1968   │     30   │  event source                                       │
  │   44592   │    186   │  thread/other                                       │
  │   36600   │     29   │  thread/spawn                                       │
  │   52344   │    228   │  untagged                                           │
  │   12232   │     59   │  wifi                                               │
  └───────────┴──────────┴─────────────────────────────────────────────────────┘
  Total: 252664 bytes in 596 allocations (98%), largest free 1k, total free 6k
I'm triggering the container every 60 seconds however the error happens more often than that
And for context- I'm working on a Crowd Tracker for my (small) college, basically just scanning BLE devices and sending the manuf data to the server for data parsing
f
I'm looking.
The looks benign. I'm starting to worry that we leak memory with
central.scan
.
That said. With BLE you never know...
a
That's what I'm thinking
right hah
f
The tls/bignum memory (~17k) seems to indicate that Artemis is in the process of connecting to the broker.
a
If this helps-
Copy code
[artemis.scheduler] INFO: job started {job: container:tracker}
Heap report @ out of memory:
  ┌───────────┬──────────┬─────────────────────────────────────────────────────┐
  │   Bytes   │  Count   │  Type                                               │
  ├───────────┼──────────┼─────────────────────────────────────────────────────┤
  │    2176   │      2   │  external byte array                                │
  │   16936   │     23   │  tls/bignum                                         │
  │   53248   │     15   │  toit processes                                     │
  │     20480 │        5 │    system    0 b476094f-1061-696e-a209-3f0ad2a0162d │
  │     24576 │        6 │    other     1 68a2c931-1e01-4f20-cfbd-6a64b500a955 │
  │      8192 │        2 │    current  12 836308cc-d8d0-e73d-6648-17a1bcd6d444 │
  │   16384   │      1   │  heap metadata                                      │
  │    4096   │      1   │  spare new-space                                    │
  │    6080   │     21   │  lwip                                               │
  │    8136   │    657   │  heap overhead                                      │
  │    1968   │     30   │  event source                                       │
  │   45584   │    206   │  thread/other                                       │
  │   36600   │     29   │  thread/spawn                                       │
  │   50536   │    225   │  untagged                                           │
  │   12232   │     59   │  wifi                                               │
  └───────────┴──────────┴─────────────────────────────────────────────────────┘
  Total: 253976 bytes in 610 allocations (99%), largest free 2k, total free 5k

******************************************************************************
Decoding by `jag`, device has version <2.0.0-alpha.142>
******************************************************************************
Allocation failed:.
  0: Session.read-handshake-message_ <sdk>/tls/session.toit:640:15
  1: Session.handshake_.<block> <sdk>/tls/session.toit:338:11
  2: Task_.with-deadline_.<block> <sdk>/core/task.toit:223:16
  3: Task_.with-deadline_      <sdk>/core/task.toit:217:3
  4: with-timeout              <sdk>/core/utils.toit:182:24
  5: with-timeout              <sdk>/core/utils.toit:165:12
  6: Session.handshake_        <sdk>/tls/session.toit:337:9
  7: Session.handshake.<block> <sdk>/tls/session.toit:281:7
  8: Session.handshake         <sdk>/tls/session.toit:227:3
  9: Socket.handshake          <sdk>/tls/socket.toit:69:14
 10: Client.try-to-reuse_.<block>.<block> <pkg:pkg-http-2.5.1>/client.toit:651:24
 11: catch.<block>             <sdk>/core/exceptions.toit:124:10
 12: catch                     <sdk>/core/exceptions.toit:122:1
 13: catch                     <sdk>/core/exceptions.toit:85:10
 14: Client.try-to-reuse_.<block> <pkg:pkg-http-2.5.1>/client.toit:646:9
 15: Client.try-to-reuse_      <pkg:pkg-http-2.5.1>/client.toit:634:3
 16: Client.web-socket_.<block> <pkg:pkg-http-2.5.1>/client.toit:367:7
 17: SmallInteger_.repeat      <sdk>/core/numbers.toit:1194:3
 18: Client.web-socket_        <pkg:pkg-http-2.5.1>/client.toit:364:19
 19: Client.web-socket         <pkg:pkg-http-2.5.1>/client.toit:336:12
 20: main.<block>              Client/tracker.toit:27:23
 21: main                      Client/tracker.toit:10:1
******************************************************************************

[artemis.scheduler] INFO: job stopped {job: container:tracker}
[artemis.synchronize] INFO: synchronized
[artemis.synchronize] INFO: synchronized
[artemis.synchronize] INFO: synchronized
f
You could try to disable Artemis-synchronization while running your program:
Copy code
import artemis

main args:
  with-timeout (Duration --m=2):
    artemis.run --offline: actual-main args

actual-main args:  // Original main
  ...
a
ill try that
f
Ah. I see now that the websocket server you connect to is a TLS server.
So the memory comes from there.
a
Yeah the production version connects to that, dev connects to an insecure server
f
Do you know where these names ("other", "current") come from?:
Copy code
│     24576 │        6 │    other     1 68a2c931-1e01-4f20-cfbd-6a64b500a955 │
  │      8192 │        2 │    current  12 836308cc-d8d0-e73d-6648-17a1bcd6d444 │
a
I don't unfortunately
f
I'm guessing 68a2c931... is the tracker.toit ?
a
Might be?
How could I check
f
good question. I'm not sure we print it when we build the pod.
a
Ah- current is the app ID
DEBUG: current state is changed {changes: { apps: +{tracker: {id: 836308cc-d8d0-e73d-6648-17a1bcd6d444, triggers: {interval: 60}, arguments: [wss://..., ...]}} }}
f
I see. So your app is using 8k. And something else is using 24.5K.
That must be Artemis.
I don't remember it being that memory hungry.
Let me check what I have.
a
yeah the system is only using 16k so its gotta be artemis
let me try running the same code without artemis
f
unfortunately I also have ~20K.
I guess we will spend a bit of time over the next weeks to trim that down a bit...
a
Here's without artemis, still crashing:
Copy code
Heap report @ out of memory:
  ┌───────────┬──────────┬─────────────────────────────────────────────────────┐
  │   Bytes   │  Count   │  Type                                               │
  ├───────────┼──────────┼─────────────────────────────────────────────────────┤
  │    1536   │      1   │  external byte array                                │
  │   36864   │     11   │  toit processes                                     │
  │     16384 │        4 │    system    0 d76b3d7e-9719-6d20-835d-fac4a92e544c │
  │     12288 │        3 │    current   1 224bbb94-c28a-2f47-fd16-27d72746f31f │
  │      8192 │        2 │    other     2 22d830b6-90fa-da4b-ef98-ea7d9a1659cf │
  │   16384   │      1   │  heap metadata                                      │
  │    4096   │      1   │  spare new-space                                    │
  │   34752   │     72   │  lwip                                               │
  │    8104   │    648   │  heap overhead                                      │
  │    2024   │     29   │  event source                                       │
  │   41752   │    153   │  thread/other                                       │
  │   31000   │     25   │  thread/spawn                                       │
  │   49992   │    230   │  untagged                                           │
  │   11952   │     56   │  wifi                                               │
  └───────────┴──────────┴─────────────────────────────────────────────────────┘
  Total: 238456 bytes in 577 allocations (93%), largest free 28k, total free 41k
f
When you say "without Artemis", how do you run the program?
This looks strange, though:
Total: 238456 bytes in 577 allocations (93%), largest free 28k, total free 41k
It looks like something wants to allocate more than 28K.
a
jag container install tracker ./Client/tracker.toit
f
Ok. So it's running the Jaguar container instead of the Artemis container.
can you do a
jag container list
?
a
Copy code
DEVICE         IMAGE                                  NAME
lucid-silver   22d830b6-90fa-da4b-ef98-ea7d9a1659cf   tracker
lucid-silver   224bbb94-c28a-2f47-fd16-27d72746f31f   jaguar
f
Ok, so it was Jaguar running out of memory (current == "224...")
I don't understand why it was trying to allocate 28K.
was there a stacktrace?
a
Copy code
[jaguar] INFO: container 'tracker' installed and started
Heap report @ out of memory:
  ┌───────────┬──────────┬─────────────────────────────────────────────────────┐
  │   Bytes   │  Count   │  Type                                               │
  ├───────────┼──────────┼─────────────────────────────────────────────────────┤
  │      56   │      2   │  tls/bignum                                         │
  │   45056   │     13   │  toit processes                                     │
  │     16384 │        4 │    system    0 d76b3d7e-9719-6d20-835d-fac4a92e544c │
  │     20480 │        5 │    other     1 224bbb94-c28a-2f47-fd16-27d72746f31f │
  │      8192 │        2 │    current   3 22d830b6-90fa-da4b-ef98-ea7d9a1659cf │
  │   16384   │      1   │  heap metadata                                      │
  │    4096   │      1   │  spare new-space                                    │
  │    6152   │     24   │  lwip                                               │
  │    7776   │    613   │  heap overhead                                      │
  │    1944   │     29   │  event source                                       │
  │   44480   │    173   │  thread/other                                       │
  │   36600   │     29   │  thread/spawn                                       │
  │   48808   │    223   │  untagged                                           │
  │   11960   │     56   │  wifi                                               │
  └───────────┴──────────┴─────────────────────────────────────────────────────┘
  Total: 223312 bytes in 549 allocations (87%), largest free 4k, total free 35k

******************************************************************************
Decoding by `jag`, device has version <2.0.0-alpha.143>
******************************************************************************
MALLOC_FAILED error.
  0: Session.handshake.<block> <sdk>/tls/session.toit:281:7
  1: Session.handshake         <sdk>/tls/session.toit:227:3
  2: Socket.handshake          <sdk>/tls/socket.toit:69:14
  3: Client.try-to-reuse_.<block>.<block> <pkg:pkg-http-2.5.1>/client.toit:651:24
  4: catch.<block>             <sdk>/core/exceptions.toit:124:10
  5: catch                     <sdk>/core/exceptions.toit:122:1
  6: catch                     <sdk>/core/exceptions.toit:85:10
  7: Client.try-to-reuse_.<block> <pkg:pkg-http-2.5.1>/client.toit:646:9
  8: Client.try-to-reuse_      <pkg:pkg-http-2.5.1>/client.toit:634:3
  9: Client.web-socket_.<block> <pkg:pkg-http-2.5.1>/client.toit:367:7
 10: SmallInteger_.repeat      <sdk>/core/numbers.toit:1194:3
 11: Client.web-socket_        <pkg:pkg-http-2.5.1>/client.toit:364:19
 12: Client.web-socket         <pkg:pkg-http-2.5.1>/client.toit:336:12
 13: main.<block>              Client/tracker.toit:27:23
 14: main                      Client/tracker.toit:10:1
******************************************************************************
f
Interesting. So "other" and "current" might not be related.
But that's not the dump from above.
a
I can see if I can recreate the issue with a local server (no TLS)
f
Is the TLS server something we can test again?
a
Right- there was no stack trace for that one
Sure
f
Some servers use certificates that are harder to work with.
a
Copy code
[jaguar] INFO: container 'tracker' installed and started
Heap report @ out of memory:
  ┌───────────┬──────────┬─────────────────────────────────────────────────────┐
  │   Bytes   │  Count   │  Type                                               │
  ├───────────┼──────────┼─────────────────────────────────────────────────────┤
  │     224   │      1   │  tls/bignum                                         │
  │   45056   │     13   │  toit processes                                     │
  │     16384 │        4 │    system    0 d76b3d7e-9719-6d20-835d-fac4a92e544c │
  │     20480 │        5 │    current   1 224bbb94-c28a-2f47-fd16-27d72746f31f │
  │      8192 │        2 │    other     4 b09117fb-db68-76de-da22-b464f69cbfd7 │
  │   16384   │      1   │  heap metadata                                      │
  │    4096   │      1   │  spare new-space                                    │
  │   29464   │     54   │  lwip                                               │
  │   12288   │   1171   │  heap overhead                                      │
  │    2024   │     29   │  event source                                       │
  │   56872   │    704   │  thread/other                                       │
  │   31008   │     25   │  thread/spawn                                       │
  │   48808   │    223   │  untagged                                           │
  │   11952   │     56   │  wifi                                               │
  └───────────┴──────────┴─────────────────────────────────────────────────────┘
  Total: 258176 bytes in 1105 allocations (101%), largest free 0k, total free 0k
Heap report @ out of memory:
  ┌───────────┬──────────┬─────────────────────────────────────────────────────┐
  │   Bytes   │  Count   │  Type                                               │
  ├───────────┼──────────┼─────────────────────────────────────────────────────┤
  │     224   │      1   │  tls/bignum                                         │
  │   45056   │     13   │  toit processes                                     │
  │     16384 │        4 │    system    0 d76b3d7e-9719-6d20-835d-fac4a92e544c │
  │     20480 │        5 │    current   1 224bbb94-c28a-2f47-fd16-27d72746f31f │
  │      8192 │        2 │    other     4 b09117fb-db68-76de-da22-b464f69cbfd7 │
  │   16384   │      1   │  heap metadata                                      │
  │    4096   │      1   │  spare new-space                                    │
  │   29464   │     54   │  lwip                                               │
  │   12360   │   1180   │  heap overhead                                      │
  │    2024   │     29   │  event source                                       │
  │   57120   │    714   │  thread/other                                       │
  │   31008   │     25   │  thread/spawn                                       │
  │   48808   │    223   │  untagged                                           │
  │   11952   │     56   │  wifi                                               │
  └───────────┴──────────┴─────────────────────────────────────────────────────┘
  Total: 258496 bytes in 1115 allocations (101%), largest free 0k, total free 0k
[jaguar] WARN: running Jaguar failed due to 'OUT_OF_MEMORY' (1/3)
I'll add some debugging logs, see if I can pinpoint where it crashes
f
It's almost certainly the ws connection.
Which server are you using?
a
Like for the websockets?
f
yes.
a
The server is using Go,
net/http
and
gorilla/websockets
for the websockets
f
Written by you?
a
Correct
f
The TLS/HTTP code was written by @erikcorry , so he is probably best for investigating this. It's currently holidays here in Denmark, so he might only be able to have a look next week. If it's possible: - could you maybe avoid the TLS connection for now - give us access next week so we can try to replicate ?
a
I'll see if theres a way I can avoid the TLS connection for now- using Cloudflare so it could be unavoidable unless I move off CF temporarily
f
hmm.
a
I'll work on getting you guys access, I'd just need to clean up the Git repo
f
cloudflare is actually one that we use a lot. We don't usually have problems with it.
Maybe it's the websocket, then.
Do you sometimes get data through, or is it always failing?
a
I'll see if I get data through
e
"other" is normally Jaguar or Artemis. Current is likely your program. System is the system process.
f
@erikcorry this is the weirdest one so far
That said, some look like they simply run out of memory.
e
It's hard to run TLS and also have bluetooth support if you don't have PSRAM. They just take a lot of space.
f
I'm doing that with Artemis all the time.
e
With the TLS you can run out of memory while doing the TLS handshake. This causes TLS to fail, and I think it can free all the memory that it was using for the handshake. You don't see the OOM before the handshake has failed, and at that point the memory has been freed that it was using.
a
Yeah without TLS it seems to work fine
Could it also just be a board limitation? Just not having enough memory 😂
I'm basically running this on some cheap ESP32 dev boards from Aliexpress
f
All ESP32 boards have the same amount of internal RAM.
e
Well having PSRAM fixes the issue, so in a way that's true. But we do aim to work on non-PSRAM devices too.
f
Some have additional external RAM.
It's a bit strange: Artemis has no problem connecting to our broker with TLS (through cloudflare). But the tracker program seems to be unable to do so (reliably).
@erikcorry does a websocket connection use more memory than a normal http connection?
a
It probably does- its a constant connection I believe
e
The initial handshake can be quite heavy, depending on the crypto primitives used and the size of the certificates.
f
AmusedGrape is using cloudflare as well.
e
After that there's a period where the memory use is not so bad, but the go servers will gradually increase the size of the encrypted blocks up to 16k. That should be manageable, and we no longer require that memory to be contiguous.
f
I think the peak usage is during connection, so the fact that the websocket stays alive probably doesn't matter too much here.
e
We can see some pretty big allocations that are not Toit.
Copy code
│   57120   │    714   │  thread/other                                       │
  │   31008   │     25   │  thread/spawn                                       │
  │   48808   │    223   │  untagged                                           │
We don't really know what these are. Our malloc implementation records them, but we don't have a lot of insight, because it's not Toit. Some of it is certainly BLE.
f
@amusedgrape you have set the max-offline to something bigger. So the tracker application starts up fresh and isn't running with a system where a BLE scan has been done before. Right?
a
Yeah max-offline is 90s
f
Ah. So there is something that could help: if you move the ble initialization after the connection to the server, the memory for the BLE won't be in use while the connection is established (where the http connection uses the most memory).
e
Yeah, the BLE grabs a lot of memory at boot time, but it probably gets even more when you start using it.
a
I'll try that
f
If that doesn't help (enough), then there is another approach: - capture the seen IDs in memory - shut down the BLE - then only establish a connection to the server. This way BLE and TLS don't need to be alive at the same time at all. It would mean that the scanned IDs are sent with a delay (of at most 20s).
a
I think I initially tried capturing the IDs in memory, but I think it captured so many that it ran out of memory
e
You can get that report at other times than OOM:
Copy code
import system show serial-print-heap-report

main:
  serial-print-heap-report
Would be interesting to see if the memory use rises when you start using BLE.
a
I'll try that
f
Maybe one more before opening BLE? (
adapter := ble.Adapter
)
a
Copy code
session := client.web-socket --uri="wss://crowdtracker.hamp.sh/api/report" --headers=headers // ws://192.168.0.58:8080/ws

  print "After WS connection, before BLE scan"
  serial-print-heap-report

  adapter := ble.Adapter
Should be how it is
f
ah. I see.
f
Weird. In the first log you gave the WS connection didn't add a lot of memory. In the last one, it used up a lot. Not sure what that could be. BLE is definitely the worst offender.
Can you send us the code that you are running again?
To me this looks like BLE + WiFi is just not a good idea. I would run BLE first. Scan for 20 seconds, and then only (after having shut down BLE) use the WiFi connection.
Since it works from time to time, it's probably just at the edge of what could be possible, though.
a
How would I shut down BLE?
f
very good question...
a
Hah, yeah I found nothing in the docs
f
It looks like we are not exposing the
close
. There is a low-level function, though. Could you try:
ble.ble-close_ adapter.resource-group_
? If that works, I will add a
close
to the
Adapter
next week.
a
Actually let me get some better heaps
Copy code
try:
    try:
      central.scan --duration=SCAN-DURATION: | device/ble.RemoteScannedDevice |
        // session.send (hex.encode device.data.manufacturer-data)
        devices.add (hex.encode device.data.manufacturer-data)
    finally:
      print "After BLE scan"
      serial-print-heap-report
      ble.ble-close_ adapter.resource-group_
      print "After BLE close"
      serial-print-heap-report
      session.send (json.encode {
        "devices": devices
      })
      print "After WS send"
      session.close
      print "After WS close"
  finally:
      network.close
f
Hmm. Could you print how many entries you have?
Fwiw, closing the BLE definitely helps.
a
804... 😅
I'll see if I can remove duplicates
f
This is (more or less), how I would try it:
Copy code
import artemis
import ble
import certificate-roots
import encoding.hex
import http
import net

SCAN-DURATION   ::= Duration --s=20

main args:
  with-timeout (Duration --m=2):
    artemis.run --offline: actual-main args

actual-main args:
  certificate-roots.install-common-trusted-roots
  adapter := ble.Adapter
  central := adapter.central

  data := {}
  
  central.scan --duration=SCAN-DURATION: | device/ble.RemoteScannedDevice |
    data.add device.data.manufacturer-data
  
  ble.ble-close_ adapter.resource_group_
  
  network := net.open
  client :=  http.Client network
  headers := http.Headers

  if args[1] == null or args[1] == "":
    throw "Authorization header is required"
    exit 1

  headers.add "Authorization" args[1]
  headers.add "Device-ID" artemis.device.id.to-string

  try:
    session := client.web-socket --uri=args[0] --headers=headers // ws://192.168.0.58:8080/ws

    data.do:
      session.send (hex.encode it)
    finally:
      session.close
  finally:
      network.close
a
14, that definitely helped lol
I'll try that code real quick
Works after some slight modification!
f
Nice.
I will add the
close
to the BLE next week.
f
Interesting. Is your container still running at that point?
a
I think? I can't really tell
I think it is?
f
Are you running with
artemis.run --offline
?
a
I disabled that- it doesn't seem to run with that for some reason
f
hmm. what does "It doesn't seem to run" mean?
What happened, was that the container did some bluetooth things while Artemis tried to synchronize. -> BLE + TLS at the same time.
And it looks like this was after a "reset", so it couldn't use TLS resume (which is a much cheaper way of establishing the connection).
a
I'll test again here shortly
Yeah seems like it is working, I just didn't know if it was running. Running offline shouldn't affect the interval trigger though right?
f
"offline" just means that Artemis is not allowed to synchronize while the block is active.
a
Ahh perfect then
f
In theory, this can make it impossible for Artemis to synchronize, which is why I added the
with-timeout
. Just to make sure Artemis actually has a chance of synchronizing in case there is a bug in the program.
a
Oh yeah that was smart
f
In theory we probably want a watchdog instead, but this was simpler
a
Kinda unrelated- what should I do in this situation:
Copy code
[artemis.synchronize] INFO: firmware update initiated
[artemis.synchronize] INFO: firmware update {from: <base64>, to: <base64>}
[artemis.synchronize] INFO: synchronized state to broker
[artemis.scheduler] INFO: runlevel decreasing {runlevel: 1}
[artemis.synchronize] INFO: firmware update {size: 1719648}
E (11105) Toit: Oversized ota_begin args: 1719648-1703936
[artemis.scheduler] INFO: runlevel increasing {runlevel: 3}
[artemis.synchronize] WARN: firmware update failed {error: OUT_OF_BOUNDS}
Would I have to update the boards in person rather than OTA?
f
I think this means that the OTA partition isn't big enough
Contrary to Jaguar, Artemis bundles the user code with the firmware. The default partition scheme thus easily runs out of space.
In your pod spec:
envelope: esp32-ota-1c0000
.
We should probably warn when users start with OTA partitions that are relatively small.
Note that you can't (easily) change the partition size over the air. So right now the easiest is to reflash.
If really necessary one can update the partition table OTA we will but it's complicated. I wrote a blog post about it: https://blog.toit.io/changing-an-esp32-partition-table-over-the-air-276c86feeba8
a
Weird, I added
envelope: esp32-ota-1c0000
to my pod spec, but getting this error when flashing:
Copy code
Firmware is too big to fit in designated partition (1719648 > 1703936)
EXCEPTION error. 
firmware tool failed with exit code 0
f
Hmm. The 1703936 is still the old size.
Did you rebuild and upload the pod?
This is happening with
artemis serial flash
?
a
Yes and yes
f
Let me try to reproduce.
While I'm trying to reproduce, could you try to read the partition table from your device? I have written a small tool that prints it: https://github.com/toitware/toit-partition-table-esp32
argh.
my mistake. It's
firmware-envelope
and not just
envelope
.
I even made the mistake in the json-schema, so the auto-completion suggests
envelope
.
I will write a patch for Artemis that allows both, and we will gradually migrate towards
envelope
. I think. (Pending review).
PR is out. If Kasper agrees, we will accept "envelope" and "firmware-envelope" in the next release. (But only document "envelope").
a
Oh here's a new one..
Broker error: duplicate key value violates unique constraint "pods_pkey"
SQL is great isn't it?
f
checking.
Which operation did you do?
a
Uploading the pod
f
From a built pod, or from a spec?
a
i think built, and i specified the file
f
ah. ok. I can reproduce.
a
./artemis pod upload --tag=v1.0.3r1 tracker.pod
Sweet
f
It looks like you uploaded the same pod again.
-> primary key error (since the pod's id is in the file).
and
artemis pod tag
is on my TODO list. Doesn't exist yet...
a
weird, somehow that was the reason
seems to all work now!
f
nice 🙂 Thanks for sending us all that feedback. My todo list has a few more items now...
a
of course! best part about beta software 😜
f
Hopefully not too much left. We are hoping to get to v2.0.0 (Toit), and v1.0.0 (Artemis) soon.
a
yall have done a great job though! made my project 10x easier.
k
For what it's worth,
artemis.run --offline
takes an optional
--timeout
argument, so you don't need to wrap it in a call to
with-timeout
.
f
Also noteworthy: since cloudflare supports TLS resume, establishing a TLS connection will be significantly cheaper if it happens within 23h of the last connection, and if the RTC memory wasn't deleted (which happens with a hard reset, or a power off, but not with deep sleep).
2 Views