This message was deleted.
# ask-for-help
s
This message was deleted.
j
cc @sauyon
s
Something that you might want to try is just accessing the database directly from your bentoml server. What's your usecase? Batch inference is being implemented but you'll have to do a similar thing in this case anyway.
m
The IO is not a big deal, I have image data in the S3 that I can access directly and the output could be done in multiple ways. What is bothering me is: • API: I can't have an HTTP request waiting for an hour (is this the proper way how to handle that in an API: https://softwareengineering.stackexchange.com/a/414356 ?) • blocking: wouldn't an hour running runner block the GPU and timeout all other real-time requests?