Kai Sun
04/17/2024, 12:03 AMdruid.processing.buffer.sizeBytes my understanding of this config is the size of intermediate buffer or merge buffer size. The question is this, if we set a small value, say 10M, will it cause groupby or topN query failure?
The underlying reason is that each thread in the processing pool when processing the hydrants/segment, it would use one piece of this buffer. Thus a large configuration of threads say 100 with buffer size as 500M would use 50G. That is pretty huge, meaning one middle manager server may hold only one or two peons.Kyle Hoondert
04/17/2024, 6:45 AMKai Sun
04/23/2024, 4:50 PMKyle Hoondert
04/24/2024, 7:37 AMintermediatePersistPeriod
Normal configuration for a peon is druid.processing.numThreads=2 with a higher number of threads sometimes configured for high query concurrency use cases.
To answer your original question about buffers - I would expect a 10MB buffer to cause query failures in many cases - 100M is the smallest seen with 300M being average for peonsKai Sun
04/24/2024, 11:30 PMintermediatePersistPeriod
> Normal configuration for a peon is druid.processing.numThreads=2 with a higher number of threads sometimes configured for high query concurrency use cases.
blocking call is probably not accurate. Just looked at the code when the peon persist a hydrant to files, it would use indexIO.loadIndex(persistedFile) to swap, which the persistedFile may be memory mapped?
The underlying reason is this -- Currently, the realtime query path for us is kind of slow in terms of latency. druid.processing.numThreads=2 would limit concurrent processed segment count to be 2. This is too small. Event we increase processing thread pool size to much large value, 40 in our case, there is still another limitation. The hydrant files in one segment are processed sequentially within the same processing thread. 20 hydrant files may mean 10s to 20s latency.Kai Sun
04/24/2024, 11:33 PM