This message was deleted.
# ask-for-help
s
This message was deleted.
šŸ 1
šŸ’¾ 1
j
cc @Sean
s
So the
request_duration
is a histogram metric. While you can calculate the average with the formula you have above, it’s much more straight forward to just use the
histogram_quantile
function. The following expression shows the 99th percentile of the ā€œiris_classifierā€ service latency.
histogram_quantile
is a very powerful PromQL function, I’d recommend that you read more into. https://prometheus.io/docs/prometheus/latest/querying/functions/#histogram_quantile
Copy code
histogram_quantile(0.99, rate(bentoml_api_server_request_duration_seconds_bucket{service_name="iris_classifier"}[1m]))
On top of using histogram quantile, you should verify that the default buckets cover the range of latency you expect from your service. By default, BentoML uses the default buckets configured by Prometheus, ranging from 0.005s to 10s. The histogram gets more accurate as the buckets gets more granular. You can update the buckets using the configuration specified here. https://docs.bentoml.org/en/latest/guides/metrics.html#request-duration
gratitude arigatou gozaimasu 2
s
Thanks @Sean for the detailed explanation
s
@Bo good candidate question for discourse.
šŸ‘ 1