This message was deleted.
# ask-for-help
s
This message was deleted.
y
My setup: • 4 models in the bentoml service ◦ 4 instances of model 1 ◦ 4 instances of model 2 ◦ 2 instances of model 3 ◦ 2 instances of model 4 ◦ all on GPUs • starting service with 20 api workers • hitting service with ~20-30 parallel workers Getting service busy error ~80% for most of the requests.
a
Is this the first request that you got?
y
you mean error messag?
a
yes
y
This is the full traceback:
Copy code
2022-11-22T20:06:26+0000 [ERROR] [api_server:2] Exception on /generate-phrases-generic [POST] (trace=a4728477755439c7a2f2756980de6d9b,span=6228acb99fadafd4,sampled=0)
Traceback (most recent call last):
  File "/disk1/projects/tzhang/cainet/venv/lib/python3.8/site-packages/bentoml/_internal/server/http_app.py", line 324, in api_func
    output = await api.func(input_data)
  File "/disk1/projects/tzhang/cainet/bentoml/bentos/cainet_model_server_prod/4ltywcdkowekscga/src/cainet_model_server_api.py", line 338, in generate_phrases_generic
    outputs = await generic_translation_runner.generate_phrases.async_run(request_data)
  File "/disk1/projects/tzhang/cainet/venv/lib/python3.8/site-packages/bentoml/_internal/runner/runner.py", line 53, in async_run
    return await self.runner._runner_handle.async_run_method(  # type: ignore
  File "/disk1/projects/tzhang/cainet/venv/lib/python3.8/site-packages/bentoml/_internal/runner/runner_handle/remote.py", line 207, in async_run_method
    raise ServiceUnavailable(body.decode()) from None
bentoml.exceptions.ServiceUnavailable: Service Busy
2022-11-22 20:06:26,383 - bentoml._internal.server.http_app - ERROR - Exception on /generate-phrases-generic [POST]
Traceback (most recent call last):
  File "/disk1/projects/tzhang/cainet/venv/lib/python3.8/site-packages/bentoml/_internal/server/http_app.py", line 324, in api_func
    output = await api.func(input_data)
  File "/disk1/projects/tzhang/cainet/bentoml/bentos/cainet_model_server_prod/4ltywcdkowekscga/src/cainet_model_server_api.py", line 338, in generate_phrases_generic
    outputs = await generic_translation_runner.generate_phrases.async_run(request_data)
  File "/disk1/projects/tzhang/cainet/venv/lib/python3.8/site-packages/bentoml/_internal/runner/runner.py", line 53, in async_run
    return await self.runner._runner_handle.async_run_method(  # type: ignore
  File "/disk1/projects/tzhang/cainet/venv/lib/python3.8/site-packages/bentoml/_internal/runner/runner_handle/remote.py", line 207, in async_run_method
    raise ServiceUnavailable(body.decode()) from None
bentoml.exceptions.ServiceUnavailable: Service Busy
a
when did you receive this err log?
y
it’s during when I’m hitting the server with ~20 parallel workers for load testing.
Here’s the runner config if you think it’s helpful:
Copy code
runners:
  model1:
    resources:
      cpu: 4
      <http://nvidia.com/gpu|nvidia.com/gpu>: [0, 0, 0, 0]
    batching:
      enabled: true
      max_batch_size: 64
      max_latency_ms: 1000
  model2:
    resources:
      cpu: 4
      <http://nvidia.com/gpu|nvidia.com/gpu>: [1, 1, 1, 1]
    batching:
      enabled: true
      max_batch_size: 64
      max_latency_ms: 1000
  model3:
    resources:
      cpu: 2
      <http://nvidia.com/gpu|nvidia.com/gpu>: [2, 2]
    batching:
      enabled: true
      max_batch_size: 32
      max_latency_ms: 100
  model4:
    resources:
      cpu: 2
      <http://nvidia.com/gpu|nvidia.com/gpu>: [2, 2]
    batching:
      enabled: true
      max_batch_size: 8
      max_latency_ms: 100
model 1&2 are slower, model 3 and 4 are faster
r
same problem, @Yilun Zhang did you find solution?
Service Busy after >150 users/requests in second
y
I was having this issue even with only 20 users or 10 users.
I’m already using async-await for API endpoints. Is there any suggestions for improving/fixing this, or at least to learn more about this?
s
I believe service unavailable means that the runner server doesn't think that it can serve your request within the max latency you set, and is therefore cancelling the request---does raising the max latency help?
y
Yeah, that’s exactly what I was thinking but i felt like with the current setup it should be raising this issue. I’m getting this even with ~10 workers. Raised latency to 10000 and still the same, but it’s not acceptable for me because the latency is too large. i also tried to host 8 models of each but still no luck with this.
s
The same as in you're still getting service unavailable errors?
y
yes, so I’m wondering if bentoml team has any baseline for # of api workers and number of models to host for certain RPS and model response time? For example, if I have a text-generation model with single request response time of ~250ms, and if i want to achieve ~30-40 RPS with ~max response time of ~500ms, then what could be the ideal setup?
s
I think unfortunately batched request time varies depending on the kind of model, so it's hard to say. @larme or @Aaron Pham might have more context, though. Ballpark I'd note that to achieve 30-40rps with that latency naively (without batching) I'd expect you to need ~15-20x the current resources to match that throughput.
y
I’m doing batching, so that’s probably the main latency (waiting to get requests to get batched). I tried to scale the resource a bit more by hosting ~20 models of each (total 4 models). Not getting service busy anymore but having really bad response time. I guess it’s moving towards better direction. I will experiment more. The mystery was, before when using 0.13, I was able to achieve good rps with only 3 GPUs and 6 workers in total, so this was my main concern and question.
s
Hm, interesting. I think @Jiang might know more about the 0.13 behavior but that's definitely concerning.
a
How large is your text generation model?
y
It's the Pegasus model (2.2Gb)