This message was deleted.
# troubleshooting
s
This message was deleted.
d
In the past I successfully used this project for non trivial sized Cassandra. Too bad the project is dead. https://github.com/Stratio/cassandra-lucene-index
p
I seem to remember someone did hook up to Lucene for GIS ---- not heard of anyone else… one moment!
d
wow, this is sick!
g
another note— you can do poor mans' full text search today with a procedure like this: 1) tokenize your free form text field during ingest: for example convert
"Here is a sentence."
to
["here", "sentence"]
(i.e. remove punctuation, split on whitespace, remove common words like
is
and
a
, ingest as array of strings). if you load this into a column named
msg_tokens
then druid will individually index each token. 2) when searching, tokenize your search string the same way, and use selector filters on
msg_tokens
(or
MV_CONTAINS
in SQL). for example: to search for
"Here is a sentence"
, use
WHERE MV_CONTAINS(msg_tokens, 'here') AND MV_CONTAINS(msg_tokens, 'sentence')
. this is super fast, since we can leverage the index for the individual tokens.
it's a bit awkward but it works!
that being said, proper full text search using something designed for that specifically would be rad
d
do you think I can perform no.1 using the embedded JS? perform tokenization and create a new column based on the result?
g
i'm reluctant to encourage anyone to enable JS for any purpose due to the security concerns. let's say you don't care about that, though. then the answer is i'm not sure since i don't remember exactly what JS is capable of 🙂. in particular i don't recall if JS transformation fns can generate arrays, or if they can only generate strings/numbers so… maybe?
i think you could almost do it with the builtin expressions: https://druid.apache.org/docs/latest/misc/math-expr.html
string_to_array
splits the string, although i think you might really want split-on-regex to really properly tokenize (remove all whitespace and punctuation) once it's an array, you can remove stop words using
map
and
replace
or, you could do the transformation upstream using Lucene's tokenizer https://lucene.apache.org/core/7_3_1/core/org/apache/lucene/analysis/standard/StandardTokenizer.html
speaking of lucene 🙂