Hi Druid Gurus, About `druid.processing.buffer.siz...
# general
k
Hi Druid Gurus, About
druid.processing.buffer.sizeBytes
my understanding of this config is the size of intermediate buffer or merge buffer size. The question is this, if we set a small value, say 10M, will it cause groupby or topN query failure? The underlying reason is that each thread in the processing pool when processing the hydrants/segment, it would use one piece of this buffer. Thus a large configuration of threads say 100 with buffer size as 500M would use 50G. That is pretty huge, meaning one middle manager server may hold only one or two peons.
k
Just for clarity here - are you saying you have 100 processing threads configured? Like you have 101 CPUs on the machine?
k
@Kyle Hoondert, sorry to reply a little bit late. Out of town for several days. No, I don't have 100 CPUs. I guess basically your point is that the processing thread count should be around the same as number of CPUs. However, in the Peon real time processing, accessing the hydrant files can be blocking calls to file systems.
k
Hi @Kai Sun - I don’t quite understand what is meant about ‘blocking calls to file system’ - peons write to the filesystem only occasionally when max rows/bytes in memory is reached, or
intermediatePersistPeriod
Normal configuration for a peon is
druid.processing.numThreads=2
with a higher number of threads sometimes configured for high query concurrency use cases. To answer your original question about buffers - I would expect a 10MB buffer to cause query failures in many cases - 100M is the smallest seen with 300M being average for peons
👍 1
k
> I don’t quite understand what is meant about ‘blocking calls to file system’ - peons write to the filesystem only occasionally when max rows/bytes in memory is reached, or
intermediatePersistPeriod
> Normal configuration for a peon is
druid.processing.numThreads=2
with a higher number of threads sometimes configured for high query concurrency use cases. blocking call is probably not accurate. Just looked at the code when the peon persist a hydrant to files, it would use
indexIO.loadIndex(persistedFile)
to swap, which the persistedFile may be memory mapped? The underlying reason is this -- Currently, the realtime query path for us is kind of slow in terms of latency.
druid.processing.numThreads=2
would limit concurrent processed segment count to be 2. This is too small. Event we increase processing thread pool size to much large value, 40 in our case, there is still another limitation. The hydrant files in one segment are processed sequentially within the same processing thread. 20 hydrant files may mean 10s to 20s latency.
To shorten the query latency, we need either reduce the number of hydrants to query, or parallel the hydrant queries. parallel the hydrant queries means that each thread of execution for one hydrant processing would need an intermediate buffer. Thus, the size of intermediate buffer is a limitation of how many concurrent hydrant queries can be performed. And thus the original question.