Slackbot
02/18/2021, 9:14 PMChaoyu
03/04/2021, 8:18 PMChaoyu
03/04/2021, 8:23 PMTerry Kong
03/04/2021, 8:33 PM2*real_time to report the result since i wait real_time to record the audio and real_time to perform the inference
“anything involving unidirectional RNNs or stateful architectures (where only part of the input is available at any given time)” - that is supported for sureCool, perhaps I missed this in the documentation. If you have pointers to something in the docs, please let me know. FWIW, i know if i accept RNN states as inputs in, say the
JsonInput , and emit the RNN states in the JsonOutput , i can effectively have a stateful transaction with a bento service, but I was curious if there was something like a cache in the bento service that we could store the states in so that the RNN states don’t need to be serialized/deserialized all the time. Perhaps this isn’t the best design choice since if we had multiple bentos behind a load balancer, there’s technically only one bento service with the correct cache that can respond to this particular query. But if you guys have already thought about this use case I’d love to hear what you guys suggest hereChaoyu
03/04/2021, 8:35 PMsave or loadChaoyu
03/04/2021, 8:38 PMChaoyu
03/04/2021, 8:40 PMChaoyu
03/04/2021, 8:41 PMPerhaps this isn’t the best design choice since if we had multiple bentos behind a load balancer, there’s technically only one bento service with the correct cache that can respond to this particular query. But if you guys have already thought about this use case I’d love to hear what you guys suggest hereI believe what you will need is sticky sessions support in the proxy layer
Terry Kong
03/04/2021, 8:45 PMYour model states or BentoService’s in-memory attributes does not get serialized or deserialized while the API server is up and running, it only happen duringPerhaps this was already clear, but the what I meant by serializing and deserializing was something like this, which I know is already possible under the bento API (allow me to write some pseudocode):orsaveload
class MyBento(...):
@api(...)
def do_inference(inputs: JsonInput):
inputs_json = inputs["inputs"]
inputs_tensor = tf.convert_to_tensor(inputs_json)
lstm_states = inputs["lstm_states"]
lstm_states_tensor = tf.convert_to_tensor(lstm_states) # "deserializing" from json -> tf.Tensor
outputs, lstm_cell_and_hidden_states = self.artifacts.model(inputs=inputs_tensor, states=lstm_states_tensor)
return {
"outputs": serialize_tensor_to_json(outputs),
"lstm_cell_and_hidden_states": serialize_tensor_to_json(lstm_cell_and_hidden_states), # extra "serialize" from tf.Tensor -> json
}
i was just wondering if bentoml had a way to avoid us having to send and receive those states because they can be quite large (tens of thousands of elements for large models)Terry Kong
03/04/2021, 8:54 PMYes gRPC is an option, it can also be done without gRPC. What’s your thoughts on concurrency & throughput? One model instance will handle one long-lived request at a time?I think one model can handle one long-lived request at a time, which is probably the simplest. Perhaps this is logic wouldn’t apply to all types of models, but multiple long lived requests can have individual “chunks” of information that could be batched, so one model could process those batched chunks and possibly return it to multiple processes that were serving concurrent long lived approaches. I haven’t thought through this completely, but naively I would expect throughput to improve since we’d now be allowed to process part of the input before the entire input has been sent. Another relevant domain here could be processing/classification of streamed video input
Terry Kong
03/04/2021, 8:58 PMBtw are you using a framework like Tensorflow or PyTorch? I don’t think they have good support for streaming input yet, and there’s very little BentoML could do if the ML runtime itself does not support itI’m using TF, and you’re right that streaming support is non existent. Even the accompanying serving framework, TF-Serving (which supports grpc), doesn’t actually support bidirectional streaming. My hopeful thought was that bentoml could be the abstraction that handles the streaming since it could maybe hold an open communication and chunk the input to a model and return partial outputs as we go