This message was deleted.
# general
s
This message was deleted.
j
Hi DK, Can you share your schema? It is possible that for some datatypes Druid may use a bit more space to store the data. Druid also currently does not use dictionaries for top level numeric columns, only for string ... numeric columns are compressed differently. If you can share your schema that would help those who know more detail on this to provide a better comparison. Thanks. John
👍 1
d
ok thanks. Here is an example of dim_product table where parquet file in S3 is 64 Mb and the size in Druid is 319 Mb.
Here is the schema for this table in druid:
product_pguid string unspsc_1 string unspsc_2 string unspsc_3 string unspsc_4 string taxonomy_id string manf_desc string prod_desc string sku string modified_date string brand string sku_manf string
j
Hmmm ... strings should all be dictionary encoded, a couple of things I know could make the overall size larger with dictionaries: • high fragmentation, i.e. lots of small segments, generating more copies of the dictionaries which might have significant overlap across segments. How big is your datasource, and what are the sizes/counts of segments? ◦ Recommendation: reindex (compact) your datasource to have larger, fewer segments if possible, to reduce the number of copies of high cardinality dictionaries being stored. • high cardinality (even unique) string fields, requiring large numbers of dictionary entries ◦ Recommendation: If there are unique fields that contain numeric values and you are not filtering on them in queries, then try storing as a numeric datatype, top level numeric columns do not use dictionaries, instead they store in compressed numeric format ◦ If there is a high cardinality field that has high overlap in values across all segments in the datasource, then you can also try configuring the datasource with Range partitioning on that column, which will shard a single copy of that dictionary values across all of the segments, instead of highly duplicated values. There is also a "front encoding" feature that may shave down the dictionary sizes significantly: https://druid.apache.org/docs/latest/ingestion/ingestion-spec#front-coding
d
thank you John. I will try to implement these suggestions