This message was deleted.
# general
s
This message was deleted.
j
Those sizes are rules of thumb ... I have seen segments up to several GB in size, and upwards of 60m rows. And in my experience the two things you want to monitor to make sure the segments aren't too large are: • Peon scan times for streaming ingestion, for when the data is still held in the streaming ingestion task. Peons are generally slower than Historicals for conducting segment scans (their data is more fragmented, and some of it is unindexed), so it can help to reduce Peon segment scan time by building smaller segments more frequently • build/publish times during ingestion when the segments are first created ... "normal" segments take 3-5 minutes to build and publish ... but larger segments can take up to 10-20 minutes or more. You just need to make sure this doesn't interfere with your ingestion task operations. • Historical segment scan times will tell you if the segments are too large. For high QPS good scan times are in the single digit ms ... for low QPS you can probably tolerate up to 300ms scan time. However the scan time is also highly dependent on the query being executed. Thanks. John
k
Thanks @John Kowtko we saw few times some queries are running in secs during that time i don't see any CPU,memory increase. trying to understand why these queries are taking time . one thing i noticed is the queries that are taking time have current timestamp which can get data from real time segments that are present on historicals .we are trying to find the root cause. is there a way to monitor the scan time on historicals and if the query is waiting on threads to process ?
to keep more data in memory . i am seeing our memory usage is around 30% . how can we increase this so that more data is in memory and queries can run faster? by increasing the heap ?
j
The metrics are all listed here: https://druid.apache.org/docs/latest/operations/metrics#historical e.g. • query/segment/time = scan time per segment • segment/scan/pending = queue length there should also be the equivalent for real-time, which indicates the "segments" scanned by the Peons to serve up queries. One thing to note about memory usage at query time ... for the Peons, the "in memory" data is not indexed so it actually can scan slower than the data that is persisted (and stored in indexed columnar format) ... so a general rule of thumb is to minimize maxRowsInMemory to get the data out of the row buffer more quickly.
k
Thanks @John Kowtko what about memory usage on Historicals ?
j
I think it's jvm/mem/max ... should give you per/JVM memory information .... sys/mem is server level ... ? I am not 100% confident because I usually look at these metrics through the Imply Clarity application, the metrics are represented under different names and grouping dimensions ...
k
ok ..