This message was deleted.
# ask-for-help
s
This message was deleted.
j
Hey @Sangeon Yong, thanks for the feedback. It seems there's quite alot going here, it would be helpful if you could help us with a minimal reproducible code for us to replicate your issue. one possible reason for the
503 Service Busy
is that you have batching turned on but only
max_latency_ms: 60
, and you mentioned your model takes a long time for inference, so that may be one possible issue here. as for the GPU memory exploding, we would love to test it out and see how we can improve on it! Would be great if you could help us with a reproducible example!
s
Thanks for your help. For me, increasing
max_latency_ms
does not solve the issue. Also, I have another question. When I set
batching: enabled
to false and set
max_batch_size
to 1, then I got the following error.
Copy code
File "/usr/local/lib/python3.11/dist-packages/starlette/routing.py", line 677, in lifespan
    async with self.lifespan_context(app) as maybe_state:
  File "/usr/lib/python3.11/contextlib.py", line 204, in __aenter__
    return await anext(self.gen)
           ^^^^^^^^^^^^^^^^^^^^^
  File "/usr/local/lib/python3.11/dist-packages/bentoml/_internal/server/base_app.py", line 75, in lifespan
    on_startup()
  File "/usr/local/lib/python3.11/dist-packages/bentoml/_internal/server/runner_app.py", line 100, in _init_metrics_wrappers
    buckets=exponential_buckets(1, 2, max_max_batch_size),
            ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/usr/local/lib/python3.11/dist-packages/bentoml/_internal/utils/metrics.py", line 44, in exponential_buckets
    assert start < end
           ^^^^^^^^^^^
AssertionError
The error does not shown when I set
max_batch_size
to 2. I understand that if I set
batching: enabled
to false, then the other settings are not used. However,
max_batch_size
still give an effect to the container. Could you explain why and how to solve the problem? Additionally, if we can reproduce the issue in the minimal reproducible code, I would share you.