Slackbot
11/17/2022, 2:16 PMJiang
11/18/2022, 3:35 AMSean
11/19/2022, 11:51 AMrequest_duration is a histogram metric. While you can calculate the average with the formula you have above, itās much more straight forward to just use the histogram_quantile function. The following expression shows the 99th percentile of the āiris_classifierā service latency. histogram_quantile is a very powerful PromQL function, Iād recommend that you read more into. https://prometheus.io/docs/prometheus/latest/querying/functions/#histogram_quantile
histogram_quantile(0.99, rate(bentoml_api_server_request_duration_seconds_bucket{service_name="iris_classifier"}[1m]))
On top of using histogram quantile, you should verify that the default buckets cover the range of latency you expect from your service. By default, BentoML uses the default buckets configured by Prometheus, ranging from 0.005s to 10s. The histogram gets more accurate as the buckets gets more granular. You can update the buckets using the configuration specified here. https://docs.bentoml.org/en/latest/guides/metrics.html#request-durationSaswat Nanda
11/19/2022, 4:18 PMSean
11/20/2022, 12:36 AM