This message was deleted.
# ask-for-help
s
This message was deleted.
c
Hi @์˜๋นˆ, assuming you are discussing multiple runners in single container deployment mode since BentoML runs each runner in its own processes, the shorter task should be able to take CPU time while the long-running task A is preempted. However, this will not be possible with GPU tasks. Another option is to utilize distributed runner deployment available in BentoCloud, where each runner can be scheduled and scaled separately in their own node groups and you can configure queuing, scaling strategy, and instance type differently for each runner, depending on their priority.
๐Ÿ‘ 1
u
@Chaoyu Thank you for the answer. It was very helpful! Have a great day.