Hey there, what is the motivation behind this? Can you help me understand your usecase?
d
Dan Hertz
09/18/2023, 5:19 PM
We have very variable sizes of items that we send to the GPU for inference, so we would want to be able to have a certain number of bytes that we consider a batch, to keep under GPU memory
a
Aaron Pham
09/18/2023, 7:02 PM
Am I correct to assume that this is a LLM use case? (Please correct if I’m wrong)
Aaron Pham
09/18/2023, 7:08 PM
This is currently not supported at the moment. What you can do is to somewhat fine tune this usage such that it meets your GPU memory requirements
d
Dan Hertz
09/18/2023, 9:06 PM
What do you mean by "this usage"? Would there be a way to open a feature request for this?
s
sauyon
09/19/2023, 2:04 AM
Feel free to open an issue on GitHub! Likely it won't be coming till the server rewrite, so it'd be a little while, unless you'd like to contribute!