hey 🙂 could you share something with me when you have some results. Sorry that I don’t have anything to share but I’m really interested in learning more
c
Chaoyu
02/17/2022, 8:27 PM
Hi @Saeid Ghafouri@joseph kobti - there isn’t an example project for this yet, but we will definitely build a gallery example project for it soon. Scaling up inference graph is actually a key feature in BentoML 1.0, attached is an example we showed in our 1.0 community demo meeting:
🙌 1
Chaoyu
02/17/2022, 8:28 PM
Under the hood, on a single node deployment, BentoML will be able to utilize all CPU cores/GPUs via multiple python processes, Runner (or each step in the inference graph) will be schedule on different processes, to run in parallel
🙌 1
Chaoyu
02/17/2022, 8:29 PM
In a distributed deployment on Kubernetes created with Yatai https://github.com/bentoml/Yatai, each step will be scheduled on their own Pods/ReplicaSet, and scale individually - this will avoid bottle neck in your inference graph pipeline, and maximize your overall resource utilization.