This message was deleted.
# ask-for-help
s
This message was deleted.
j
1. Performance wise on the architecture level should be the same, since it is both BentoML Runner under the hood. The difference will be the implementation details of the runner. 2. I think batching is not available on OpenLLM's model afaik. 3. Yes @Aaron Pham CMIIW
a
2. Batching is disabled for all OpenLLM runners 3. OpenLLM has option to mask directly to this multi-GPU assignment
--workers-per-resource
. Note that by default openllm has a custom strategies and will only create 1 runner instance at all time
1. OpenLLM Runners are more specialized comparing to
bentoml.transformers
runners. It is designed for specific model architecture
y
Thanks @Jian Shen Yap and @Aaron Pham for the quick response! Sorry I got pulled away and didn’t finish the reply. So from your answers, if my end goal is to 1) host a LLM on a machine with multiple GPUs with bentoml, 2) enable dynamic batching, my option is to go with BentoML and create a custom runner class to define my own generation method, very much the same as hosting a standard transformers model, right? For the configuration and device mapping, I will just use the standard bentoml configuration yaml file.
a
Sure, imho batching doesn’t help a lot when working with LLM that requires more than 1 GPU to load in. If you have multiple A100s then I think you should write your own custom Runner and define your own generation method. I think if you want to run on smaller GPU or multiple GPUs that is not A10g and A100, having more than one instance of runner doesn’t help anyway, since you will probably need to use all of the GPUs in order to load the model
y
We do have a number of A10/100 GPUs, so each model is loaded on a single GPU, and potentially multiple instances on the same GPU, so that’s not really an issue. We are experimenting to maximize throughput right now. For example, with 1 40GB A100 GPU, I should be able to load 2 7b models instances in bfloat16,
a
yes
I think you can just use
openllm.Runner
in this instance
Copy code
runner = openllm.Runner("llama", model_id="meta-llama/Llama-2-7b-chat-hf")
Then in the config, this runner can be set with
Copy code
runners:
  llm-llama-runner:
    resources:
      <http://nvidia.com/gpu|nvidia.com/gpu>: [0]
Or you can always go the route of implementing your own generation method
y
Thanks! I do want to have more control in the prompt construction by creating extra methods in the runner class, since the input itself isn’t full prompt but rather pieces of placeholders in the prompt.
a
I think this can leave to the client to do this
prompt construction shouldn’t be at the server level
but it makes sense if you want to include the prompt into the server as well
I’m actually working on improving the client for openllm a bit to support prompt construction 😄
y
Cool, looking forward to more updates! But I think I will proceed with the older BentoML format for now so that I have better control
a
sounds good
y
Thank you for the help!
a
Happy to help! Please let me know if there are anything I can help with
🙌 1