This message was deleted.
# ask-for-help
s
This message was deleted.
j
Copy code
2023-09-14T15:12:58+0000 [ERROR] [api_server:2] Exception on /score_items [POST] (trace=6f48876ae962ed514c1d46631defd14c,span=ac35b83c971d4b35,sampled=0,service.name=item_scorer)
Traceback (most recent call last):
  File "/usr/local/lib/python3.11/site-packages/bentoml/_internal/server/http_app.py", line 341, in api_func
    output = await api.func(*args)
             ^^^^^^^^^^^^^^^^^^^^^
  File "/home/bentoml/bento/src/service.py", line 168, in score_items
    return await runner.predict.async_run(inp)
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/usr/local/lib/python3.11/site-packages/bentoml/_internal/runner/runner.py", line 55, in async_run
    return await self.runner._runner_handle.async_run_method(self, *args, **kwargs)
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/usr/local/lib/python3.11/site-packages/bentoml/_internal/runner/runner_handle/remote.py", line 248, in async_run_method
    raise ServiceUnavailable(body.decode()) from None
bentoml.exceptions.ServiceUnavailable: Service Busy
This error isn't very informative - any idea why my runner is 'Busy'?
This should 1000% be documented
Is there no way to enable batching but not drop large quantities of requests?
j
What would be your expected behaviour Judah? because normally even for simple server that couldn't handle requests within the timeout limit, requests would get dropped
j
I'd expect the batching latency to be a target and for the timeout to be respected if the batching latency can't be met
If batching is disabled the the --timeout and --backlog parameters determine if a request is rejected, right?
If that's the case then I don't see why that couldn't still be the case when batching is enabled
The max_latency_ms then becomes an upper bound to hold the request queue before dispatching a batch (or when the max_batch_size is hit)
s
Honestly at the moment our timeout / SLO controls are super convoluted.
--timeout
is actually the client-side timeout, and
max_latency_ms
is what we consider to be the (runner) server-side timeout. Really we should get #3630 in for more fine-grained control but it's been pending on us to come up with some sort of benchmark for a while now...
👏 1
j
#3630 looks great! Is there anything I can do to help it along?
Looks like this is a common point of confusion atm (3x Slack threads on it in the last month) so it seems like it should probably be prioritized given that it is also core functionality?
j
I'll bring it up to the team and give you some updates of our future actionable!
🫶 2
t
Thank you. I fixed the error after increase max_latency_ms.
Hi, everyone. The response time of my API is around 13 seconds. I set max_latency_ms=86400000 (1 day) but still got error 503 Server was busy after sending 100 requests in 5 minutes. How can I fix that? Thanks!
j
I ended up just disabling batching 😞
👀 1
t
In my case, doing that would lead to worse results.
j
hey @THE AI do you mind running the serve with
--debug
and see if you could share us mroe informative logs? The team is currently putting cycles into looking into this
t
My error was gone after I increased runners.traffic.timeout (I missed this config cause a similar argument api_server.traffic.timeout - in my case this one for 504 Gateway Timeout), I believe increases runners.traffic.timeout, and runners.batching.max_latency_ms could solve 503 Service Busy. Thank you @Jian Shen Yap @Judah Rand @sauyon for your help.
j
Glad that you found your answer! the team will be working on bubbling up better error message for the adaptive batching, instead of just showing
503 Service Busy
. Thanks for your feedbacks so far!
🫶 1