Question about bloom filter. In the latest doc (<h...
# pinot-dev
j
Question about bloom filter. In the latest doc (https://docs.pinot.apache.org/basics/indexing/bloom-filter), I see that the following statement
Copy code
Support for raw value columns is WIP.
is no longer there. Does this mean bloom filter is now available on raw forward index columns w/o the need to enable dictionary? My use-case is to enable bloom filter on a mostly random integer column.
g
Yes, AFAIK encoding is now no important when using bloom filters
j
Got it. Is there any future planned work to allow bloom-filter to be available on a per-compressed-column-chunk basis or finer grained than per-segement based? We have a table column where the values are basically random data. Enabling dictionary encoding or text-index (after casting it to string) on it would be too expensive in terms of memory + storage footprint. So we currently use raw-forward-index + brute-force scan. We wanted to explore and see if there are some option which would allow us to reduce the brute-force scan over potentially TBs of data in a single column for every query.
g
What I wanted to say is that documentation may be not updated. I think we support bloom filters in raw columns right now.
👋 1
j
Are there any future planned work related to making bloom filter operate on a more fine-grained chunks of data?