This message was deleted.
# ask-for-help
s
This message was deleted.
👀 1
🍱 1
b
Hello @Thomas Jacquemin are you running in k8s cluster? You can typical monitoring the system usage that way.
t
Yes, the bentoml server will be running in k8s cluster. I'm asking because I see that other ML inference server solutions tend to propose such metrics by default on their
/metrics
endpoint. (ex: Triton Inference Server or TorchServe for instance)
For CPU and memory system metrics, I can rely on k8s monitoring. But what about GPU metrics consumed by each bento runner ? I don't think I can have the detailed utilization for runners natively with k8s.
b
I see. For now, we recommend to use native k8 cluster to monitoring system metrics. As for GPU usage, I would recommend you to look into nvidia’s DCGM.
👍 1
t
Thank you for your suggestion ! I'll look into it :)