Slackbot
10/10/2023, 6:22 AMJian Shen Yap
10/10/2023, 10:48 AM503 Service Busy is that you have batching turned on but only max_latency_ms: 60 , and you mentioned your model takes a long time for inference, so that may be one possible issue here.
as for the GPU memory exploding, we would love to test it out and see how we can improve on it! Would be great if you could help us with a reproducible example!Sangeon Yong
10/10/2023, 11:50 AMmax_latency_ms does not solve the issue.
Also, I have another question. When I set batching: enabled to false and set max_batch_size to 1, then I got the following error.
File "/usr/local/lib/python3.11/dist-packages/starlette/routing.py", line 677, in lifespan
async with self.lifespan_context(app) as maybe_state:
File "/usr/lib/python3.11/contextlib.py", line 204, in __aenter__
return await anext(self.gen)
^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib/python3.11/dist-packages/bentoml/_internal/server/base_app.py", line 75, in lifespan
on_startup()
File "/usr/local/lib/python3.11/dist-packages/bentoml/_internal/server/runner_app.py", line 100, in _init_metrics_wrappers
buckets=exponential_buckets(1, 2, max_max_batch_size),
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib/python3.11/dist-packages/bentoml/_internal/utils/metrics.py", line 44, in exponential_buckets
assert start < end
^^^^^^^^^^^
AssertionError
The error does not shown when I set max_batch_size to 2. I understand that if I set batching: enabled to false, then the other settings are not used. However, max_batch_size still give an effect to the container. Could you explain why and how to solve the problem?
Additionally, if we can reproduce the issue in the minimal reproducible code, I would share you.