This message was deleted.
# ask-for-help
s
This message was deleted.
a
Hey there, what is the motivation behind this? Can you help me understand your usecase?
d
We have very variable sizes of items that we send to the GPU for inference, so we would want to be able to have a certain number of bytes that we consider a batch, to keep under GPU memory
a
Am I correct to assume that this is a LLM use case? (Please correct if I’m wrong)
This is currently not supported at the moment. What you can do is to somewhat fine tune this usage such that it meets your GPU memory requirements
d
What do you mean by "this usage"? Would there be a way to open a feature request for this?
s
Feel free to open an issue on GitHub! Likely it won't be coming till the server rewrite, so it'd be a little while, unless you'd like to contribute!