Slackbot
11/11/2022, 5:57 AMDidip Kerabat
11/11/2022, 5:59 AMPeter Marshall
11/11/2022, 7:31 AMPeter Marshall
11/11/2022, 7:32 AMDidip Kerabat
11/12/2022, 3:54 AMGian Merlino
11/15/2022, 5:47 AM"Here is a sentence." to ["here", "sentence"] (i.e. remove punctuation, split on whitespace, remove common words like is and a, ingest as array of strings). if you load this into a column named msg_tokens then druid will individually index each token.
2) when searching, tokenize your search string the same way, and use selector filters on msg_tokens (or MV_CONTAINS in SQL). for example: to search for "Here is a sentence", use WHERE MV_CONTAINS(msg_tokens, 'here') AND MV_CONTAINS(msg_tokens, 'sentence'). this is super fast, since we can leverage the index for the individual tokens.Gian Merlino
11/15/2022, 5:49 AMGian Merlino
11/15/2022, 5:50 AMDidip Kerabat
11/15/2022, 5:50 AMGian Merlino
11/15/2022, 6:11 AMGian Merlino
11/15/2022, 6:13 AMGian Merlino
11/15/2022, 6:15 AMstring_to_array splits the string, although i think you might really want split-on-regex to really properly tokenize (remove all whitespace and punctuation)
once it's an array, you can remove stop words using map and replaceGian Merlino
11/15/2022, 6:15 AMGian Merlino
11/15/2022, 6:16 AM