Hmmm ... strings should all be dictionary encoded, a couple of things I know could make the overall size larger with dictionaries:
• high fragmentation, i.e. lots of small segments, generating more copies of the dictionaries which might have significant overlap across segments. How big is your datasource, and what are the sizes/counts of segments?
◦ Recommendation: reindex (compact) your datasource to have larger, fewer segments if possible, to reduce the number of copies of high cardinality dictionaries being stored.
• high cardinality (even unique) string fields, requiring large numbers of dictionary entries
◦ Recommendation: If there are unique fields that contain numeric values and you are not filtering on them in queries, then try storing as a numeric datatype, top level numeric columns do not use dictionaries, instead they store in compressed numeric format
◦ If there is a high cardinality field that has high overlap in values across all segments in the datasource, then you can also try configuring the datasource with Range partitioning on that column, which will shard a single copy of that dictionary values across all of the segments, instead of highly duplicated values.
There is also a "front encoding" feature that may shave down the dictionary sizes significantly:
https://druid.apache.org/docs/latest/ingestion/ingestion-spec#front-coding