This message was deleted.
# ask-for-help
s
This message was deleted.
e
I think this may fall into the category of "offline" inference. I've heard that bento is working on a new offline inference capability, but I don't think it's out yet
f
Thanks! I think it's not properly "offline" inference. "Offline" inference means inference not in a real-time context (not in a user-facing app), like when you want to debug or during a batch job. In my case, the worker would be deployed in a real-time context, it's just that it would receive samples to predict from a queue instead of HTTP request. I'll make an issue on the GitHub, I think it could be a nice feature to have 🙂
e
Sure. I think I'd seen bento use that term to describe: any time you run inference not in REST API mode.
Here are the docs of the previous version of Bento I'd found on that: https://docs.bentoml.org/en/0.13-lts/guides/batch_serving.html
This seems potentially useful, but you have to invoke bento as a CLI rather than somehow call it directly from code which is disappointing to me. I'm hoping hte new version in bento 1.0.0+ allows for direct API calls in code
b
@Eric Riddoch With the recently released bentoml client, you should able to call directly in python environment
e
Is it out?? Do you have docs I can read on it? Super excited
b
Here you go, fresh out of the printer(1.0.10) https://docs.bentoml.org/en/latest/guides/client.html
e
Nice! Is this the recommended approach for batch inference now? If so, I'd love to experiment with it. We use Metaflow for our batch jobs. If we can invoke this on large batches of data in our Metaflow DAGs... Bento will probably end up everywhere we run inference.