This message was deleted.
# troubleshooting
s
This message was deleted.
d
LIKE
on our largest clusters is like a foot-bazooka that can blow up at anytime.
j
do you allow user send sql directly to druid? or you can do some filter before sending to druid?
d
The queries come from Superset mostly.
We can probably add custom logic on Superset side.
p
In the recesses of my mind I seem to have a memory of @Gian Merlino a few years back saying something about LIKE %x and LIKE x% having different performance characteristics… but I have drunk a lot of coffee since then :P
g
yeah
LIKE 'x%'
reads a range of the index without checking individual values, and
LIKE '%x'
or
LIKE '%x%'
reads the entire index and checks the pattern against every distinct value
so on high-cardinality columns,
LIKE 'x%'
is substantially faster than
LIKE '%x'
or
LIKE '%x%'
on low-card columns they're all about the same
(all should be fast)
d
ok, that’s similar to MySQL, however none of these use-cases are using the bitmap index, right?
r
I think that the like operation is ran on the dictionary encoded (that's why it's faster/safe to stop scanning then the like is on the start) and then, the logic to use or not the bitmap index may vary depending on the number of results, if any if the number of results is too high, druid may spending too much time during bitmaps AND's and a seq-scan could be better in some cases (but also, which struct could be good to check the existence of the matched values? another map lookup or a b-tree with the values IDs? ) but I don't know if the druid even do any different matching execution depending on the number of matches
g
they would all use the bitmap indexes
although on high-card columns that still involves, for non-prefix
LIKE
, applying the
LIKE
pattern to each distinct value in the dictionary (so we know which bitmaps to pull)
it can still take some time for those high-card columns, even though we're using an index